Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
arXiv:2608.05780v1 Announce Type: new Abstract: Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selectin