Research
When More Words Say Less: Decoupling Length and Specificity in Image Description Evaluation
arXiv:2601.04609v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used to make visual content accessible via text-based descriptions. In current systems, however, desc
arXiv:2601.04609v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used to make visual content accessible via text-based descriptions. In current systems, however, description specificity is often conflated with their length. We argue that these two concepts must be disentangled: descriptions can be concise yet dense with information, or lengthy yet vacuous. We define specificity relative to a contrast set, where a description is more specific to the extent that it picks out the target image better than other possible images. We construct a dataset that controls for length while varying information content, and validate that people reliably prefer more specific descriptions regardless of length. We find that controlling for length alone cannot account for differences in specificity: how the length budget is allocated makes a difference. These results support evaluation approaches that directly prioritize specificity over verbosity.
Related
- VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
- Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
- Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts
- Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models
Source: arXiv cs.CL | 2026-04-21