Applications
Big unsaturated benchmarks that have this: ARC-AGI, the original GDPval (not GDPval-AA), METR long horizons, ASI cyber tasks, (Speaking of w…
Ethan Mollick highlights that large, currently under‑explored benchmarks (e.g., ARC‑AGI, GDPval, METR long horizons, ASI cyber tasks) are growing in complexity, yet they increasingly lack systematic h
Ethan Mollick highlights that large, currently under‑explored benchmarks (e.g., ARC‑AGI, GDPval, METR long horizons, ASI cyber tasks) are growing in complexity, yet they increasingly lack systematic human performance baselines. He argues that validated AI benchmarks must include multiple human references to remain meaningful, noting the rising difficulty and expense of acquiring such data as tests expand. This underscores a critical challenge in frontier AI evaluation: balancing advanced task design with reliable human comparison metrics.
Related
- Annoying that OpenAI doesn’t seem to give a GDPval measure for GPT 5.6. One of the best measures of economically valuable work.
- Thresholds (like “can a robot make a good coffee?”), are not actually very good benchmarks because you can’t track progress towards a goal.
- It would also be useful for funding R&D into benchmarking models, which is currently mostly done by the labs themselves right now.
Source: Ethan Mollick (X) | 2026-07-30