Model Releases
CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models
arXiv:2608.20763v1 Announce Type: cross Abstract: Linear probes and activation steering have uncovered that vision-language models (VLMs) internally represent mental states such as agents' beliefs, kn
arXiv:2608.20763v1 Announce Type: cross Abstract: Linear probes and activation steering have uncovered that vision-language models (VLMs) internally represent mental states such as agents' beliefs, knowledge, and intentions. However, it is unclear whether and how these representations are used by downstream predictions along these axes. To close this gap, we introduce Cross-Axis Routing Diagnostic (CARD), which steers activations along one axis while measuring the response of a different axis's prediction. Applied to open-weight VLMs on Relay Chain -- a new cooperative grid-world benchmark we propose -- we diagnose a critical routing failure: models fail to incorporate belief representations into their next action prediction, effectively leaving valuable information about their partners unused.
Related
- Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models
- Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
- Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
Source: arXiv cs.AI | 2026-08-24