Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
DGX agentarXiv:2607.24407v1 Announce Type: new Abstract: Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and comp