Model Releases
On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5 scores 63.6% with a 69.6% p…
On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5 scores 63.6% with a 69.6% pass rate on Extended — approaching Fable 5 at half the cost.
On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5 scores 63.6% with a 69.6% pass rate on Extended — approaching Fable 5 at half the cost. Within Devin, it shows particular strength on difficult debugging and root-cause analysis tasks. https://devin.ai/blog/claude-opus-5
Related
- On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability and quality, Sonnet 5 scores 53.8% and h…
- Claude Fable 5 is now available in Devin. Fable 5 earns the #1 spot on FrontierCode, our benchmark for real-world engineering tasks that gra…
- Right alongside Cursor, Devin Desktop (Windsurf) and CLI now support GLM-5.2 as well. FrontierCode Extended is a benchmark we care deeply ab…
- GPT-5.6 is now available in Devin! On FontierCode 1.1 Extended, the GPT-5.6 family stands out for pairing strong scores with excellent cost …
Source: Cognition AI (X) | 2026-07-24