Local Ai

Comprehension of Multilingual Expressions Referring to Target Objects in Visual Inputs

arXiv:2511.11427v2 Announce Type: replace Abstract: Referring Expression Comprehension (REC) requires models to localize objects in images based on different types of natural language descriptions. Ev

DGX agentpaper
local-aiarxiv-cs-cv

arXiv:2511.11427v2 Announce Type: replace Abstract: Referring Expression Comprehension (REC) requires models to localize objects in images based on different types of natural language descriptions. Even with significant progress, research on the area remains predominantly English-centric, despite increasing global deployment demands. This work addresses multilingual REC through two main contributions. First, we construct a unified multilingual dataset spanning 10 languages, by systematically expanding 12 existing English REC benchmarks through machine translation and context-based translation enhancement. Second, we introduce an attention-anchored efficient neural architecture that uses a multilingual SigLIP2 encoder. Our attention-based approach generates coarse spatial anchors from attention distributions, which are subsequently refined through learned residuals. Experimental evaluation demonstrates competitive performance on standard benchmarks despite the use of a relatively small model. Multilingual evaluation shows consistent capabilities across languages, establishing the practical feasibility of efficient multilingual visual grounding systems.

Source: arXiv cs.CV | 2026-08-18

Loading related sources…