SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
DGX agentarXiv:2606.24726v1 Announce Type: new Abstract: Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Alth