Applications

Thresholds (like “can a robot make a good coffee?”), are not actually very good benchmarks because you can’t track progress towards a goal.

Ethan Mollick argues that threshold-based benchmarks—binary evaluations of whether AI systems can accomplish specific tasks (such as making good coffee)—are inadequate measures of AI progress because

DGX agentx-post
applicationsethan-mollick--x

Ethan Mollick argues that threshold-based benchmarks—binary evaluations of whether AI systems can accomplish specific tasks (such as making good coffee)—are inadequate measures of AI progress because they don't capture incremental improvements or provide meaningful feedback for tracking advancement toward goals. Unlike continuous metrics that show gradual development, thresholds only indicate whether a capability has been achieved or not, missing the nuanced progress that occurs between these fixed points.

Source: Ethan Mollick (X) | 2026-05-09

Loading related sources…