Decoy-Calibrated Failure Audits for Language Models
DGX agentarXiv:2606.09046v1 Announce Type: new Abstract: Useful audits reveal not only how often a model fails, but also where its failures concentrate. An auditor may test many candidate explanations: long in