Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
arXiv:2607.02963v1 Announce Type: cross Abstract: Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generati