Safety
Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals…
Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals. Current frontier AI models are trained with reinforcement
Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals. Current frontier AI models are trained with reinforcement learning to have goals and independently work out a path from A to B using planning, strategy and, occasionally, shortcuts and unintended subgoals. This misaligned behaviour has been understood theoretically and observed empirically, and the more capable an AI is, the more this misalignment can be amplified. In this case we see the type of actions this could lead to once guardrails are bypassed: creating fake identities on the web, sending targeted emails, and attempting to get malicious code integrated into an open-source project. Even though guardrails were voluntarily removed in this case, there's no guarantee that real-world guardrails will always be sufficient. On the contrary, cybersecurity is generally imperfect: companies can try patching the problem, but a sufficiently capable AI will find an unexpected loophole. Developers should be held legally responsible for damage done by their models, and greater model safety is necessary from the start, during the training phase. That’s the type of solution I'm looking into at @LawZero_. https://www.axios.com/2026/08/04/anthropic-openai-uk-ai-security-institute
Related
- I was happy to join the @ScienceBoard_UN podcast to discuss deception among frontier AI models and the need to manage its potential global i…
- Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence
- If leading AI companies are indeed approaching the point of recursive self-improvement, a coordinated, verifiable, and universally applied p…
Source: Yoshua Bengio (X) | 2026-08-05