Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
DGX agentarXiv:2510.08480v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text