Research

In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tosse…

In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tossed it completely For maximum transparency we just open source

DGX agentx-post
researchfrancois-chollet--x

In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tossed it completely For maximum transparency we just open sourced all the human replays for public games My favorite charts that come out of these are the Action Progression charts. It's much easier to get a feel for how efficient environment solves turn into scores ARC-AGI-3 Human Baseline Dataset Today we're open-sourcing the ARC-AGI-3 Human Baseline. This is the most exhaustive human testing study in the ARC-AGI series Every environment was solved by at least 2 people (many by more) from the general public, with no prior training

Related

Source: Francois Chollet (X) | 2026-04-14

Loading related sources…