Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
DGX agentarXiv:2607.08066v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misalign