Research
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
arXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the toke
arXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks (p {<} 0.01, McNemar). (iii) Oracle analysis reveals model complementarity up to +27.8pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain (+1.6 to +5.7pp) but are domain-specific and hurt transfer (-1.2pp). These results argue for budget-conditioned evaluation protocols.
Source: arXiv cs.AI | 2026-08-13