A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth
DGX agentarXiv:2601.21817v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm.