Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
DGX agentarXiv:2605.27101v1 Announce Type: cross Abstract: A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actua