Research
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
arXiv:2507.13868v2 Announce Type: replace Abstract: Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their interna
arXiv:2507.13868v2 Announce Type: replace Abstract: Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual input can lead to hallucinations and unreliable predictions. In this work, we investigate the mechanisms that VLMs use to resolve cross-modal conflicts by introducing WHOOPS-AHA!, a dataset of multimodal counterfactual queries that deliberately contradict internal commonsense knowledge. Through logit inspection, we identify a small set of attention heads that mediate this conflict. By intervening in these heads, we can steer the model towards its internal parametric knowledge or the visual information. Our results show that attention patterns on these heads effectively locate image regions that influence visual overrides, providing a more precise attribution compared to gradient-based methods.
Related
- Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
- Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding
- VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
- Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
Source: arXiv cs.CV | 2026-04-21