Safety
Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build …
Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build upon them and use them to evaluate the monitorability of the
Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build upon them and use them to evaluate the monitorability of their own models. https://alignment.openai.com/monitorability-evals/
Related
- Anthropic details using AI agents to accelerate alignment research on 'weak-to-strong supervision', where a weak model supervises the training of a stronger one (Anthropic)
- Benchmarking Misuse Mitigation Against Covert Adversaries
- Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4
- Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence
Source: Sam Altman (X) | 2026-04-24