Constitutional Black-Box Monitoring for Scheming in LLM Agents
arXiv:2603.00829v2 Announce Type: replace-cross Abstract: Safe deployment of Large Language Model (LLM) agents in autonomous settings requires reliable oversight mechanisms. A central challenge is det