iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
DGX agentarXiv:2605.31096v1 Announce Type: new Abstract: While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language model