Model Releases

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-eval…

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics' I really thought it would treat 'now do benc

DGX agentx-post
model-releasesethan-mollick--x

Ha! It did it: "We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics" I really thought it would treat "now do benchbenchbenchbenchbench" as a joke, but Sol actually did reasonable experiments. As a joke I prompted Codex "Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a good arXiv paper." I got a PDF. But the paper is actually kind of interesting?

Related

Source: Ethan Mollick (X) | 2026-07-25

Loading related sources…