Model Releases

The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of m…

The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of money developing a very good benchmark of autonomous model ab

DGX agentx-post
model-releasesethan-mollick--x

The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of money developing a very good benchmark of autonomous model ability at hard tasks, GDPval, and haven't reported it for GPT-5.6. We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a…

Source: Ethan Mollick (X) | 2026-07-09

Loading related sources…