The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
DGX agentarXiv:2603.22278v2 Announce Type: replace Abstract: Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their prop