The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
DGX agentarXiv:2606.30653v1 Announce Type: cross Abstract: Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification