AI Evaluation Should Measure Verification Cost, Not Correctness Alone
DGX agentarXiv:2608.08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to verify those o