Research
To score 100% on a game, you just need to beat the median action efficiency of an unfiltered pool of random people. Easy if you're a bit sma…
Francois Chollet makes the observation that achieving a perfect score on certain AI benchmark games or tasks only requires surpassing the median performance of an unfiltered random population sample,
Francois Chollet makes the observation that achieving a perfect score on certain AI benchmark games or tasks only requires surpassing the median performance of an unfiltered random population sample, rather than demonstrating exceptional intelligence or skill. This framing suggests that some benchmarks marketed as challenging may have a relatively low bar, since beating the median of random, untrained participants is achievable for anyone slightly above average. The post likely continues with a critique of how such benchmarks are designed or interpreted as measures of intelligence.
Related
- In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tosse…
- Any smart human giving it real effort should score >90% on ARC-AGI-3
- ARC-AGI-3 Human Baseline Dataset Today we're open-sourcing the ARC-AGI-3 Human Baseline. This is the most exhaustive human testing study in …
- Simply retrieving a reasoning trace looks a lot like human reasoning, until it's time to navigate uncharted territory. If you memorized all …
Source: Francois Chollet (X) | 2026-04-15