Research
MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
arXiv:2604.16313v1 Announce Type: cross Abstract: Retrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal struct
arXiv:2604.16313v1 Announce Type: cross Abstract: Retrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal structures. While retrieval-augmented generation (RAG) has shown strong performance in text-based QA, its extensions to multimodal documents remain underexplored and face significant limitations. Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. MARA consists of two components: a Query-Aligned Region Encoder that builds multi-level document representations and reweights them based on query relevance to improve retrieval precision; and a Self-Reflective Evidence Controller that monitors evidence sufficiency during generation and adaptively incorporates content from lower-ranked sources using a sliding-window strategy. Experiments on six multimodal QA benchmarks demonstrate that MARA consistently improves retrieval relevance and answer quality over existing SOTA method.
Related
- Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
- MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
- Rag Performance Prediction for Question Answering
- Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
- Indexing Multimodal Language Models for Large-scale Image Retrieval
- How Retrieved Context Shapes Internal Representations in RAG
Source: arXiv cs.CL | 2026-04-21