Model Releases
A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and co…
A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and compacting context lets the model build on what it has already
A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and compacting context lets the model build on what it has already learned. https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
Source: OpenAI (X) | 2026-07-29