Model Releases
GuideGround: VLM-guided Semantic Understanding and Viewpoint-aware Reasoning for 3D Visual Grounding
arXiv:2608.00518v1 Announce Type: new Abstract: 3D visual grounding aims to localize the target object in a 3D scene from a natural language query, requiring both fine-grained semantic understanding a
arXiv:2608.00518v1 Announce Type: new Abstract: 3D visual grounding aims to localize the target object in a 3D scene from a natural language query, requiring both fine-grained semantic understanding and viewpoint-dependent spatial reasoning. Existing methods typically formulate semantic understanding as an auxiliary closed-set object classification task and rely on multi-view feature aggregation for viewpoint reasoning, limiting semantic generalization and weakening viewpoint-specific evidence. We observe that vision-language models naturally provide complementary capabilities through open-vocabulary semantic understanding and global scene perception. Based on this insight, we propose GuideGround, a VLM-guided framework that complements rather than replaces task-specific grounding models by leveraging VLMs for semantic enhancement and viewpoint-specific hypothesis verification. Specifically, we replace auxiliary closed-set object classification with VLM-generated object semantic descriptions to enhance semantic understanding. Meanwhile, instead of directly aggregating multi-view representations, we preserve viewpoint-specific grounding hypotheses through per-view grounding and explicitly verify them using VLMs across candidate viewpoints. Extensive experiments on the ReferIt3D benchmark demonstrate that GuideGround consistently outperforms previous state-of-the-art methods. Comprehensive ablation studies further confirm the effectiveness of both the proposed semantic understanding and viewpoint reasoning strategies.
Related
- GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning
- GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision
- Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
- SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
Source: arXiv cs.CV | 2026-08-04