CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
arXiv:2605.18621v1 Announce Type: cross Abstract: Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, vi