Research
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
arXiv:2512.02791v2 Announce Type: replace Abstract: Dialogue-Based Generalized Referring Expression Comprehension (GREC) requires models to ground the expression and unlimited targets in complex visua
arXiv:2512.02791v2 Announce Type: replace Abstract: Dialogue-Based Generalized Referring Expression Comprehension (GREC) requires models to ground the expression and unlimited targets in complex visual scenes while resolving coreference across a long dialogue context. However, existing systems struggle under distribution shift between training and evaluation domains, a gap exacerbated by the scarcity of annotated dialogue grounding data. We address this challenge with a three-tier data-synthesis method that balances realism and controllability to produce scalable supervision for dialogue-conditioned grounding. Fine-tuning on the synthesized data yields consistent, substantial improvements over prior approaches across standard evaluation metrics.
Related
- From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality
- MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
- AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
- MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
Source: arXiv cs.CL | 2026-04-28