Model Releases
Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models
arXiv:2608.20975v1 Announce Type: new Abstract: Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channe
arXiv:2608.20975v1 Announce Type: new Abstract: Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.
Related
- CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models
- Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
- Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
- OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
Source: arXiv cs.AI | 2026-08-24