When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
DGX agentarXiv:2608.03918v1 Announce Type: cross Abstract: Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence.