Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
arXiv:2605.29800v1 Announce Type: new Abstract: LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a frame