Model Releases
// There Is No Neutral Harness // Great work discussing some of the issues in harness evaluation. Twelve open-weight models answer the same …
// There Is No Neutral Harness // Great work discussing some of the issues in harness evaluation. Twelve open-weight models answer the same 3,679 items from ARC, HellaSwag, MMLU, and TruthfulQA under
// There Is No Neutral Harness // Great work discussing some of the issues in harness evaluation. Twelve open-weight models answer the same 3,679 items from ARC, HellaSwag, MMLU, and TruthfulQA under 26 equally defensible harness configurations. Items, weights, and greedy decoding stay fixed. Only option order, prompt wording, and whether the answer is read from generated text or per-option likelihoods change. gemma4-31b lands anywhere from 31% to 89% depending on the harness alone. On the items that two adjacent models both answer stably, the pair is tied. Config-fragile items carry 95.7% of the gap between them, and four of the twelve models reach rank one under some configuration. Item discrimination, the property benchmark-compression methods maximize when picking a representative subset, correlates with fragility. Compressed benchmarks are selecting for the items most sensitive to configuration. Paper: https://arxiv.org/abs/2608.21382 Track more trending AI papers in our academy: https://academy.dair.ai/
Related
- // The Harness Effect // (bookmark it) Now more that ever pay very close attention to the orchestration harness and its effect on costs and …
- New open-source agent harness just landed! I got early access to TrueForge by TrueFoundry and have been running it locally for the past few …
- Great paper on why agent leaderboard comparisons are hard to trust. It's on the hot topic of how much of an agent benchmark score actually b…
Source: DAIR.AI (X) | 2026-08-25