Model Releases
MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs
arXiv:2607.08970v2 Announce Type: replace-cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observa
arXiv:2607.08970v2 Announce Type: replace-cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, which uses active viewpoint selection and evidence fusion to improve four base models by 12.3--20.0 percentage points under a six-image cap matching the fixed-view baseline; budget-extended gains are model-dependent and reach 27 percentage points for GPT-5.
Related
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
- Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
- How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study
- Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI
Source: arXiv cs.AI | 2026-08-10