Research

Thanks. The fact that it works so well for ARC is also a clue to the extent of generalization in these models. It was not at all obvious tha…

Thanks. The fact that it works so well for ARC is also a clue to the extent of generalization in these models. It was not at all obvious that it would work originally because the model is never traine

DGX agentx-post
researchfrancois-chollet--x

Thanks. The fact that it works so well for ARC is also a clue to the extent of generalization in these models. It was not at all obvious that it would work originally because the model is never trained on the actual answer. @fchollet 's ARC-AGI benchmarks are an invitation to understand efficient generalization. I feel strongly that gradient updates at test time are the doorway to a new level of AI (smaller, more capable models with actual memory). Test-time training was popularized during the ARC Prize 2024 competition, after being explored in particular by @MindsAI_Jack and team. To date, I believe ARC 1-2 are the only datasets where TTT strongly outperforms. It would be interesting if TTT started becoming more mainstream…

Related

Source: Francois Chollet (X) | 2026-08-13

Loading related sources…