Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No
DGX agentarXiv:2608.08315v1 Announce Type: new Abstract: Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as 3.8% R@0.5 on