Catching One in Five: LLM-as-Judge Blind Spots in Production Multi-Turn Transaction Agents
DGX agentarXiv:2606.10315v1 Announce Type: cross Abstract: LLM-as-judge is the default instrument for evaluating conversational agents, yet its reliability is almost always reported as agreement with human rat