Applications

We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choic…

We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choices about API settings, harness design, and prompting. If you

DGX agentx-post
applicationsopenai--x

We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choices about API settings, harness design, and prompting. If you’re an API developer trying to maximize performance, we recommend using the same settings that we deploy in our own products: - Use our Responses API, not our legacy Chat - Completions API - Retain reasoning - Use compaction If you want to test your own mettle against frontier models, try the public games yourself at http://arcprize.org/tasks

Source: OpenAI (X) | 2026-07-29

Loading related sources…