Research
Where To Look? : Causal Tracing of Vision Encoders in VLM
arXiv:2608.10758v1 Announce Type: new Abstract: Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actua
arXiv:2608.10758v1 Announce Type: new Abstract: Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.
Related
- From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
- Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
- Understanding How MLLMs Describe Artworks Using Token Activation Maps
Source: arXiv cs.CV | 2026-08-12