CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models
arXiv:2605.20165v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect gen