VC-Tooler: Learning Compositional and Adaptive Visual Tool Use
DGX agentarXiv:2608.02217v1 Announce Type: new Abstract: Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool int