Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
DGX agentarXiv:2607.11844v2 Announce Type: replace Abstract: Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos inv