Ask the Right Comparison:Bias-Aware Bayesian Active Top-k Ranking with LLM Judges
DGX agentarXiv:2607.02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models