PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
DGX agentarXiv:2605.23883v1 Announce Type: cross Abstract: Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this wo