Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety
DGX agentarXiv:2606.06529v1 Announce Type: new Abstract: An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework f