Research
Thanks. The fact that it works so well for ARC is also a clue to the extent of generalization in these models. It was not at all obvious tha…
Thanks. The fact that it works so well for ARC is also a clue to the extent of generalization in these models. It was not at all obvious that it would work originally because the model is never traine
Thanks. The fact that it works so well for ARC is also a clue to the extent of generalization in these models. It was not at all obvious that it would work originally because the model is never trained on the actual answer. @fchollet 's ARC-AGI benchmarks are an invitation to understand efficient generalization. I feel strongly that gradient updates at test time are the doorway to a new level of AI (smaller, more capable models with actual memory). Test-time training was popularized during the ARC Prize 2024 competition, after being explored in particular by @MindsAI_Jack and team. To date, I believe ARC 1-2 are the only datasets where TTT strongly outperforms. It would be interesting if TTT started becoming more mainstream…
Related
- I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a 'harness'), running at inference …
- If you want to help the world make sense of AGI and accelerate its arrival, consider joining the ARC Prize foundation. Two roles currently o…
Source: Francois Chollet (X) | 2026-08-13