Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
DGX agentarXiv:2512.10359v1 Announce Type: cross Abstract: Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive, understand,