Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
DGX agentarXiv:2605.24216v1 Announce Type: cross Abstract: Monitoring autonomous large language model (LLM) agents for covert malicious behavior is challenging due to delayed, context-dependent, and long-horiz