Model Releases
As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchb…
As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a
As a joke I prompted Codex "Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a good arXiv paper." I got a PDF. But the paper is actually kind of interesting?
Related
- The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of m…
- An easy way to get a team engaged with AI is just to build the thing you are talking about in the meeting during the meeting using Codex or …
- AI reviewers then ranked the submissions, and gave the same ordering every time, regardless of model doing the ranking: Codex GPT-5.4 > GPT-…
- And now a new DeepSeek model, and appears to be fully open weights. Good benchmarks, but with open models, that isn't always as meaningful. …
Source: Ethan Mollick (X) | 2026-07-24