Research
DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering
arXiv:2607.23921v1 Announce Type: new Abstract: Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an
arXiv:2607.23921v1 Announce Type: new Abstract: Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an image, which has garnered significant attention due to its practical applications. In this paper, we introduce the Dynamic Dual-level Vision Transformer Fusion Network (DDVT) for answer grounding in visual question answering. Specifically, we propose a question-guided dynamic regional-level module (QGDR) that combines complementary image context through ROI Align and text content, enabling precise localization of text-related visual content. Moreover, we present a cross-modal multi-scale aggregation module (CMA) that enhances feature fusion between pixel-level and region-level features, facilitating the effective localization of visual content associated with grounded answers. Furthermore, we fuse the located visual content with text features to locate the region and provide answers to questions posed about the image. Experimental results demonstrate that our DDVT outperforms state-of-the-art methods on several widely-used benchmarks.
Related
- Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting
- KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
- Enhancing Part-Level Point Grounding for Any Open-Source MLLMs
- Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation
- Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
Source: arXiv cs.CV | 2026-07-28