Model Releases
RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that …
RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that it is performing a completely different task it was trained
RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that it is performing a completely different task it was trained on GPT-5.5 Scores .43% on ARC AGI 3! - GPT-5.5: 0.43% - Opus 4.7: 0.18% - GPT-5.4: 0.20% - Claude 4.6: 0.45% - Gemini 3.1: 0.4% The reported failures for GPT 5.5 were: - True local effect, false world model - Wrong level of abstraction from training data - Solved the level, didn’t r…
Source: Francois Chollet (X) | 2026-05-01