Model Releases

Oh my God! @METR_Evals’s coding benchmarks are saturated! 🤯 Mythos broke the METR graph 🤯 4 weeks later, out comes a new coding task, this…

Oh my God! @METR_Evals’s coding benchmarks are saturated! 🤯 Mythos broke the METR graph 🤯 4 weeks later, out comes a new coding task, this time from @cognition: “FrontierCode Diamond remains unsaturat

DGX agentx-post
model-releasesgary-marcus--x

Oh my God! @METR_Evals’s coding benchmarks are saturated! 🤯 Mythos broke the METR graph 🤯 4 weeks later, out comes a new coding task, this time from @cognition: “FrontierCode Diamond remains unsaturated: the best performing model, Claude Opus 4.8, achieves a score of only 13.4%. There is still a lots of headroom. *Note that METR itself never panicked. It’s the Twitterverse that has egg on its face. Introducing FrontierCode: a coding eval that raises the bar for difficulty & quality. Each task took 40+ hrs of work by leading open-source maintainers. Models write sloppy code that works but isn’t maintainable. Our eval is first to measure: would you actually merge this code?

Source: Gary Marcus (X) | 2026-06-08

Loading related sources…