Research
Any smart human giving it real effort should score >90% on ARC-AGI-3
ARC-AGI-3 is a new benchmark iteration designed so that any intelligent human applying genuine effort should achieve greater than 90% accuracy, maintaining the benchmark's core principle that tasks mu
ARC-AGI-3 is a new benchmark iteration designed so that any intelligent human applying genuine effort should achieve greater than 90% accuracy, maintaining the benchmark's core principle that tasks must be solvable by humans without specialized knowledge. Francois Chollet, the creator of the ARC benchmark series, made this claim to emphasize that the test measures general fluid intelligence rather than acquired expertise or memorized knowledge. This human-accessibility threshold is a defining characteristic of the ARC benchmark family, ensuring that any gap between human and AI performance reflects a genuine difference in reasoning ability rather than domain-specific training.
Related
- ARC-AGI-3 Human Baseline Dataset Today we're open-sourcing the ARC-AGI-3 Human Baseline. This is the most exhaustive human testing study in …
- To score 100% on a game, you just need to beat the median action efficiency of an unfiltered pool of random people. Easy if you're a bit sma…
- The role of memorization and knowledge is to cache & reuse past cognitive work. It should be leveraged as a way to speed up cognition, not a…
- Simply retrieving a reasoning trace looks a lot like human reasoning, until it's time to navigate uncharted territory. If you memorized all …
Source: Francois Chollet (X) | 2026-04-15