Safety
LLM-as-Judge on a Budget
arXiv:2602.15481v2 Announce Type: replace Abstract: LLM-as-a-judge has emerged as a cornerstone technique for evaluating large language models by leveraging LLM reasoning to score prompt-response pair
arXiv:2602.15481v2 Announce Type: replace Abstract: LLM-as-a-judge has emerged as a cornerstone technique for evaluating large language models by leveraging LLM reasoning to score prompt-response pairs. Since LLM judgments are stochastic, practitioners commonly query each pair multiple times to estimate mean scores accurately. This raises a critical challenge: given a fixed computational budget B, how to optimally allocate queries across K prompt-response pairs to minimize estimation error? We present a principled variance-adaptive approach leveraging multi-armed bandit theory and concentration inequalities. Our method dynamically allocates queries based on estimated score variances, concentrating resources where uncertainty is highest. Further, our algorithm is shown to achieve a worst-case score-estimation error of ilde{O}left(sqrt{frac{sum_{i=1}^K sigma_i^2}{B}}right), sigma_i^2 being the unknown score variance for pair i in [K] with near-optimal budget allocation. Experiments on Summarize-From-Feedback and HelpSteer2 demonstrate that our method significantly outperforms uniform allocation, reducing worst-case estimation error while maintaining identical budgets. Our work establishes a theoretical foundation for efficient LLM evaluation with practical implications for AI safety, model alignment, and automated assessment at scale.
Related
- Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
- Large Reasoning Models Learn Better Alignment from Flawed Thinking
- Reinforcement-aware Knowledge Distillation for LLM Reasoning
- Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs
Source: arXiv cs.LG | 2026-04-14