Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling
DGX agentarXiv:2607.01830v1 Announce Type: new Abstract: Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a