No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
DGX agentarXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as bus