Model Releases
The latest crop of models remains below 1% on ARC-AGI-3 -- for now. Where will the scores be by the end of the year?
The latest crop of models remains below 1% on ARC-AGI-3 -- for now. Where will the scores be by the end of the year? GPT-5.5 & Opus 4.7 on ARC-AGI-3 - GPT-5.5: 0.43% - Opus 4.7: 0.18% We found 3 failu
The latest crop of models remains below 1% on ARC-AGI-3 -- for now. Where will the scores be by the end of the year? GPT-5.5 & Opus 4.7 on ARC-AGI-3 - GPT-5.5: 0.43% - Opus 4.7: 0.18% We found 3 failure modes: - True local effect, false world model - Wrong level of abstraction from training data - Solved the level, didn’t reinforce the reward See our full analysis 🧵
Source: Francois Chollet (X) | 2026-05-01