Model Releases

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures …

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures whether the agent reached it the permitted way, which is the

DGX agentx-post
model-releasesdair-ai--x

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures whether the agent reached it the permitted way, which is the question enterprise deployments care about. If you ship skills files, policy documents, or long system prompts, you have been trusting that they actually bind agent behavior. But how are you measuring all of this? Surge AI built a benchmark to actually check this. HANDBOOK.md places a standard operating procedure of 20 to 124 pages in context and grades whether it governed every action across an extended tool-use horizon. 65 tasks, five domains, ten fictional companies. Each task runs in a self-contained company environment with a file workspace plus mock email, chat, calendar, issue-tracking, and commerce services exposed over MCP. Every task mutates one of ten base handbooks, altering the specific rules and thresholds that grading turns on, so memorization does not help. Grading is fully deterministic and two-sided. 824 programmatic criteria check that required actions occurred and that prohibited actions did not. Paper: https://arxiv.org/abs/2607.25398 Learn to build effective AI agents in our academy: https://academy.dair.ai/

Source: DAIR.AI (X) | 2026-07-29

Loading related sources…