(Towards) Scalable Reliable Automated Evaluation with Large Language Models
DGX agentarXiv:2607.28282v1 Announce Type: new Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated