Applications
As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to…
As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans Validated benchmarks need to have human (ideally mul
As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans Validated benchmarks need to have human (ideally multiple humans) baselines. It is increasingly hard & pricey to do, but important
Related
- New paper (on an old AI) tests o1 against doctors on medical benchmarks & real ER cases: “across a variety of scenarios and applications, th…
- I think Epoch does a great job benchmarking, but I continue to believe that open weights models are much more fragile, especially out-of-dis…
- This post is a real Voight-Kampff test for bots on X.
Source: Ethan Mollick (X) | 2026-07-30