Safety
If alignment issues are becoming big enough that OpenAI is willing to commit 20% of research inference compute to chain-of-thought monitorin…
If alignment issues are becoming big enough that OpenAI is willing to commit 20% of research inference compute to chain-of-thought monitoring, that suggests that alignment issues are becoming a pretty
If alignment issues are becoming big enough that OpenAI is willing to commit 20% of research inference compute to chain-of-thought monitoring, that suggests that alignment issues are becoming a pretty serious concern. We really need universal policies & standards across labs. We temporarily slowed some frontier training to strengthen security and monitoring. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations help us test safeguards and gather more evidence of alignment. I expect confidence in safety to inc…
Related
- Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish
- Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
- Evading Chain-of-Thought Monitoring Through Model Poisoning
Source: Ethan Mollick (X) | 2026-08-18