Masked Diffusion Vision-Language Models for Temporal Action Localization
DGX agentarXiv:2605.29858v1 Announce Type: new Abstract: Temporal action localization (TAL) requires recognizing the target event and localizing its start and end times precisely in untrimmed videos. Recent vi