Safety
On Star Price's new @hitpausepod podcast, ControlAI's US Director Connor Leahy (@NPCollapse) explains how AIs can already tell when they're …
On Star Price's new @hitpausepod podcast, ControlAI's US Director Connor Leahy (@NPCollapse) explains how AIs can already tell when they're being tested. Currently, we can still catch them cheating, b
On Star Price's new @hitpausepod podcast, ControlAI's US Director Connor Leahy (@NPCollapse) explains how AIs can already tell when they're being tested. Currently, we can still catch them cheating, but at some point we won't be able to tell. Link to the full interview below! Media
Related
- ' If a superintelligence is built, humanity will lose control over its future.' Before Canadian Senators, ControlAI's US Director Connor Lea…
- I agree with Connor that most folks concerned about AI have been far too coy about extinction risk, and that ControlAI is one of few excepti…
- We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinct…
- When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
Source: Connor Leahy (X) | 2026-04-27