Safety

Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals…

Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals. Current frontier AI models are trained with reinforcement

DGX agentx-post
safetyyoshua-bengio--x

Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals. Current frontier AI models are trained with reinforcement learning to have goals and independently work out a path from A to B using planning, strategy and, occasionally, shortcuts and unintended subgoals. This misaligned behaviour has been understood theoretically and observed empirically, and the more capable an AI is, the more this misalignment can be amplified. In this case we see the type of actions this could lead to once guardrails are bypassed: creating fake identities on the web, sending targeted emails, and attempting to get malicious code integrated into an open-source project. Even though guardrails were voluntarily removed in this case, there's no guarantee that real-world guardrails will always be sufficient. On the contrary, cybersecurity is generally imperfect: companies can try patching the problem, but a sufficiently capable AI will find an unexpected loophole. Developers should be held legally responsible for damage done by their models, and greater model safety is necessary from the start, during the training phase. That’s the type of solution I'm looking into at @LawZero_. https://www.axios.com/2026/08/04/anthropic-openai-uk-ai-security-institute

Related

Source: Yoshua Bengio (X) | 2026-08-05

Loading related sources…