Model Releases
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3: 1. Not okay: harnesses that were custom-made to solve the b…
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3: 1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark forma
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3: 1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark format / contents. 2. Fine: general-purpose API settings that were not developed for ARC-AGI-3 and that are available to all API users. In the past, we've had a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction. I'm glad they're starting to figure out the answer. Of course, if each provider uses different settings when getting their model tested, it creates a potential parity issue. My take is that this is fine as long as the settings and the cost are clearly reported.
Source: Francois Chollet (X) | 2026-07-30