PEEK: Picking Essential frames via Efficient Knowledge distillation
DGX agentarXiv:2605.31029v1 Announce Type: new Abstract: Video-language models can process only a limited number of frames, making frame selection a key bottleneck for efficient video captioning. Most captioni