Safety
As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforceme…
As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for
As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. https://openai.com/index/pacing-model-development-cyber-capabilities/
Related
- Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss.…
- Exploration Hacking: Can LLMs Learn to Resist RL Training?
- ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment
Source: OpenAI (X) | 2026-08-18