Logit Lens Supervision for Patch-Level Explanations in Vision-Language Models
arXiv:2602.01530v2 Announce Type: replace Abstract: Modern autoregressive Vision-Language Models (VLMs) can generate fluent answers while their visual-token representations become weakly tied to the i