Model Releases

// There Is No Neutral Harness // Great work discussing some of the issues in harness evaluation. Twelve open-weight models answer the same …

// There Is No Neutral Harness // Great work discussing some of the issues in harness evaluation. Twelve open-weight models answer the same 3,679 items from ARC, HellaSwag, MMLU, and TruthfulQA under

DGX agentx-post
model-releasesdair-ai--x

// There Is No Neutral Harness // Great work discussing some of the issues in harness evaluation. Twelve open-weight models answer the same 3,679 items from ARC, HellaSwag, MMLU, and TruthfulQA under 26 equally defensible harness configurations. Items, weights, and greedy decoding stay fixed. Only option order, prompt wording, and whether the answer is read from generated text or per-option likelihoods change. gemma4-31b lands anywhere from 31% to 89% depending on the harness alone. On the items that two adjacent models both answer stably, the pair is tied. Config-fragile items carry 95.7% of the gap between them, and four of the twelve models reach rank one under some configuration. Item discrimination, the property benchmark-compression methods maximize when picking a representative subset, correlates with fragility. Compressed benchmarks are selecting for the items most sensitive to configuration. Paper: https://arxiv.org/abs/2608.21382 Track more trending AI papers in our academy: https://academy.dair.ai/

Related

Source: DAIR.AI (X) | 2026-08-25

Loading related sources…