Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
DGX agentarXiv:2607.06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicio