Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
DGX agentarXiv:2512.08410v2 Announce Type: replace Abstract: Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose