Model Releases
Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
arXiv:2608.21170v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction be
arXiv:2608.21170v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentation of a task shapes model performance and failure modes when the underlying reasoning problem is unchanged. We study this question in SPaRC, a benchmark for grid-based visual spatial planning, by introducing lightweight input-side scaffolds that preserve the visual modality while making spatial structure more accessible. Across multiple VLMs, these scaffolds improve task accuracy over the original visual setting by up to 34.0 percentage points and further complement GRPO-based training, yielding up to 4.6 additional accuracy points compared with near-zero gains on the original visual input. Analyses on both end-to-end task solving and object detection show that these gains are closely tied to reductions in grounding-related errors, while rule reasoning remains comparatively challenging. We find that visual presentation is a central factor that determines whether VLM benchmarks measure grounded perception, downstream reasoning, or a mixture of both.
Related
- Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI
- SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
- ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models
Source: arXiv cs.AI | 2026-08-24