Model Releases
We conducted cyber evaluations of Claude Mythos Preview and found that it is the first model to complete an AISI cyber range end-to-end. 🧵
Boris Cherny and colleagues conducted cybersecurity evaluations of Claude Mythos Preview, finding it to be the first AI model to successfully complete an AISI (AI Safety Institute) cyber range end-to-
Boris Cherny and colleagues conducted cybersecurity evaluations of Claude Mythos Preview, finding it to be the first AI model to successfully complete an AISI (AI Safety Institute) cyber range end-to-end. This represents a notable benchmark in AI cyber capability evaluations, as completing a full cyber range end-to-end demonstrates advanced autonomous offensive or defensive security task completion. The finding likely has significant implications for AI safety and capability thresholds tracked by the AISI.
Related
- So the concern over Mythos and cybersecurity seems warranted.
- UK gov's Mythos AI tests help separate cybersecurity threat from hype
- Mythos is very powerful, and should feel terrifying. I am proud of our approach to responsibly preview it with cyber defenders, rather than …
- Cybersecurity analysis: Claude Mythos Preview had a 73% success rate on expert-level capture-the-flag challenges, which no model could finish before April 2025 (AI Security Institute)
- Anthropic's Project Glasswing - restricting Claude Mythos to security researchers - sounds necessary to me
Source: Boris Cherny (X) | 2026-04-13