Applications

Big unsaturated benchmarks that have this: ARC-AGI, the original GDPval (not GDPval-AA), METR long horizons, ASI cyber tasks, (Speaking of w…

Ethan Mollick highlights that large, currently under‑explored benchmarks (e.g., ARC‑AGI, GDPval, METR long horizons, ASI cyber tasks) are growing in complexity, yet they increasingly lack systematic h

DGX agentx-post
applicationsethan-mollick--x

Ethan Mollick highlights that large, currently under‑explored benchmarks (e.g., ARC‑AGI, GDPval, METR long horizons, ASI cyber tasks) are growing in complexity, yet they increasingly lack systematic human performance baselines. He argues that validated AI benchmarks must include multiple human references to remain meaningful, noting the rising difficulty and expense of acquiring such data as tests expand. This underscores a critical challenge in frontier AI evaluation: balancing advanced task design with reliable human comparison metrics.

Related

Source: Ethan Mollick (X) | 2026-07-30

Loading related sources…