Model Releases
The efficiency frontier! Where do you think GPT-5.6 will land?
The efficiency frontier! Where do you think GPT-5.6 will land? Claude Opus 4.8 has landed on DeepSWE Bench, posting a 58% Pass@1 and taking #2 overall behind GPT-5.5. It continues a broader trend: sli
The efficiency frontier! Where do you think GPT-5.6 will land? Claude Opus 4.8 has landed on DeepSWE Bench, posting a 58% Pass@1 and taking #2 overall behind GPT-5.5. It continues a broader trend: slightly behind on raw score, but among the most reliable and efficient coding models across recent benchmarks.
Source: DAIR.AI (X) | 2026-05-30