Model Releases
CVT-Bench: Probing Spatial-State Integrity through Counterfactual Viewpoint Transformations
arXiv:2603.21114v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) perform strongly on isolated spatial tasks, but whether their predictions remain persistent and mutually co
arXiv:2603.21114v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) perform strongly on isolated spatial tasks, but whether their predictions remain persistent and mutually coherent across viewpoints and competing scenes is unclear. We formalize this behavioral property as spatial-state integrity and introduce CVT-Bench, a factorial diagnostic suite spanning two domains (CVT-Synthetic, CVT-Real), four context regimes (CVT-Isolated, irrelevant filler, attribute filler, CVT-Competing), three representations (Image, Text/BBox, Scene Graph), and ten viewpoint conditions (nine azimuths from 0^irc to 360^irc at 45^irc increments, plus Top) with target views withheld. Across 200 scenes and 11,833 relational queries, we measure counterfactual accuracy, cycle consistency, the normalized survival coefficient, spatial realizability, and context-specific interference. Five state-of-the-art MLLMs exhibit rapid persistence loss and frequently produce jointly unrealizable states despite high local accuracy, with failures amplified by competing contexts and natural-scene complexity. Matched irrelevant and attribute filler controls show that prompt length, position, or loss of scene access alone are insufficient explanations. On average, Text/BBox improves accuracy, persistence, and realizability, while Scene Graph provides complementary gains; neither eliminates the instability. Thus, isolated spatial accuracy substantially overestimates robustness, establishing spatial-state integrity as a distinct evaluation target. The complete benchmark and codebase will be publicly released.
Related
- AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration
- SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
- Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
- OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning
- Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations
Source: arXiv cs.CV | 2026-08-17