Model Releases

Gemini 3 Pro was the first model to achieve at least 23% on ARC-AGI-2, which it did in November, 2025 (it actually scored 31%). So the 8-12 …

Gemini 3 Pro was the first model to achieve at least 23% on ARC-AGI-2, which it did in November, 2025 (it actually scored 31%). So the 8-12 month gap between closed and open weights models still seems

DGX agentx-post
model-releasesethan-mollick--x

Gemini 3 Pro was the first model to achieve at least 23% on ARC-AGI-2, which it did in November, 2025 (it actually scored 31%). So the 8-12 month gap between closed and open weights models still seems to hold. But they are also more jagged, better at some tasks, worse at others. GLM-5.2 from @Zai_org on ARC-AGI (Verified) - ARC-AGI-2: 22.8%, 0.25 - ARC-AGI-1: 77.0%, 0.19 Performance is comparable with GPT-5.4 & 5.5 (Low Reasoning Effort)

Source: Ethan Mollick (X) | 2026-06-24

Loading related sources…