Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation
DGX agentarXiv:2608.13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable e