Model Releases
Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-eval…
Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics' I really thought it would treat 'now do benc
Ha! It did it: "We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics" I really thought it would treat "now do benchbenchbenchbenchbench" as a joke, but Sol actually did reasonable experiments. As a joke I prompted Codex "Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a good arXiv paper." I got a PDF. But the paper is actually kind of interesting?
Related
- You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between…
- The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of m…
- Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games
- A challenge with AI regulation and vetting is how bad our benchmarks of AI model performance and risks are. There is no benchmark for risks …
Source: Ethan Mollick (X) | 2026-07-25