Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
DGX agentarXiv:2605.01630v1 Announce Type: new Abstract: Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scorin