Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
arXiv:2605.12517v1 Announce Type: cross Abstract: Vision-language models (VLMs) are often deployed on text-only inputs, although they are trained with images. We find that removing the vision modality