Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
DGX agentarXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the toke