Safety
Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
arXiv:2604.13731v1 Announce Type: new Abstract: Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existin
arXiv:2604.13731v1 Announce Type: new Abstract: Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existing OCR-free methods face a trade-off between capacity and precision: end-to-end models scale poorly with document length, while visual retrieval-based pipelines are brittle and passive. We propose Doc-V^, an extbf{OCR-free agentic} framework that casts multi-page DocVQA as sequential evidence aggregation. Doc-V^ begins with a thumbnail overview, then actively navigates via semantic retrieval and targeted page fetching, and aggregates evidence in a structured working memory for grounded reasoning. Trained by imitation learning from expert trajectories and further optimized with Group Relative Policy Optimization, Doc-V^* balances answer accuracy with evidence-seeking efficiency. Across five benchmarks, Doc-V^* outperforms open-source baselines and approaches proprietary models, improving out-of-domain performance by up to extbf{47.9%} over RAG baseline. Other results reveal effective evidence aggregation with selective attention, not increased input pages.
Related
- MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed Bandits
- MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
- OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
- Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
Source: arXiv cs.CL | 2026-04-16