Model Releases

Picking the right agent harness is now a crucial skill for any AI engineer. Imagine using the same model, same task, and same prompt. Now mo…

Picking the right agent harness is now a crucial skill for any AI engineer. Imagine using the same model, same task, and same prompt. Now move it between two agent harnesses and the cost per success c

DGX agentx-post
model-releasesdair-ai--x

Picking the right agent harness is now a crucial skill for any AI engineer. Imagine using the same model, same task, and same prompt. Now move it between two agent harnesses and the cost per success can swing by 5 to 30x. This benchmark measured this across six large reasoning models, two real harnesses, 24 deterministic coding tasks with hidden evaluators, and 4,643 valid runs. Asking a model to develop and compare several approaches raised reasoning tokens by 2.4 to 7.4x with no correctness gain. Generic think-deeply cues added another 1.6 to 2.2x. A bounded-efficiency template that specifies scope, acceptance criteria, and a stop condition came out cost-neutral and sometimes halved reasoning. Harness design and prompt wording decide most agent spend before the model reasons at all, and both are cheap to change. Paper: https://arxiv.org/abs/2608.01347 Track more trending AI papers in our academy: https://academy.dair.ai/

Source: DAIR.AI (X) | 2026-08-04

Loading related sources…