Model Releases

RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that …

RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that it is performing a completely different task it was trained

DGX agentx-post
model-releasesfrancois-chollet--x

RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that it is performing a completely different task it was trained on GPT-5.5 Scores .43% on ARC AGI 3! - GPT-5.5: 0.43% - Opus 4.7: 0.18% - GPT-5.4: 0.20% - Claude 4.6: 0.45% - Gemini 3.1: 0.4% The reported failures for GPT 5.5 were: - True local effect, false world model - Wrong level of abstraction from training data - Solved the level, didn’t r…

Source: Francois Chollet (X) | 2026-05-01

Loading related sources…