Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
DGX agentarXiv:2603.08091v2 Announce Type: replace Abstract: Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by j