Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
DGX agentarXiv:2605.07106v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made remarkable progress on vision-language reasoning, yet most methods still compress visual evidence int