Research
In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tosse…
In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tossed it completely For maximum transparency we just open source
In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tossed it completely For maximum transparency we just open sourced all the human replays for public games My favorite charts that come out of these are the Action Progression charts. It's much easier to get a feel for how efficient environment solves turn into scores ARC-AGI-3 Human Baseline Dataset Today we're open-sourcing the ARC-AGI-3 Human Baseline. This is the most exhaustive human testing study in the ARC-AGI series Every environment was solved by at least 2 people (many by more) from the general public, with no prior training
Related
- ARC-AGI-3 Human Baseline Dataset Today we're open-sourcing the ARC-AGI-3 Human Baseline. This is the most exhaustive human testing study in …
- Any smart human giving it real effort should score >90% on ARC-AGI-3
- To score 100% on a game, you just need to beat the median action efficiency of an unfiltered pool of random people. Easy if you're a bit sma…
- Simply retrieving a reasoning trace looks a lot like human reasoning, until it's time to navigate uncharted territory. If you memorized all …
Source: Francois Chollet (X) | 2026-04-14