VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
DGX agentarXiv:2604.21396v1 Announce Type: cross Abstract: The advancement of Large Vision-Language Models (LVLMs) requires precise local region-based reasoning that faithfully grounds the model's logic in act