Model Releases

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning

arXiv:2604.23198v1 Announce Type: new Abstract: Current video moment retrieval excels at action-centric tasks but struggles with narrative content. Models can see extit{what is happening} but fail to

DGX agentpaper
model-releasesarxiv-cs-ai

arXiv:2604.23198v1 Announce Type: new Abstract: Current video moment retrieval excels at action-centric tasks but struggles with narrative content. Models can see extit{what is happening} but fail to reason extit{why it matters}. This semantic gap stems from the lack of extbf{Theory of Mind (ToM)}: the cognitive ability to infer implicit intentions, mental states, and narrative causality from surface-level observations. We introduce extbf{StoryTR}, the first video moment retrieval benchmark requiring ToM reasoning, comprising 8.1k samples from narrative short-form videos (shorts/reels). These videos present an ideal testbed. Their high information density encodes meaning through subtle multimodal cues. For instance, a glance paired with a sigh carries entirely different semantics than the glance alone. Yet multimodal perception alone is insufficient; ToM is required to decode that a character smiling'' may actually be concealing hostility.'' To teach models this reasoning capability, we propose an extbf{Agentic Data Pipeline} that generates training data with explicit three-tier ToM chains (intent decoding, narrative reasoning, boundary localization). Experiments reveal the severity of the reasoning gap: Gemini-3.0-Pro achieves only 0.53 Avg IoU on StoryTR. However, our 7B extbf{Shorts-Moment} model, trained on ToM-guided data, improves +15.1% relative IoU over baselines, demonstrating that extit{narrative reasoning capability matters more than parameter scale}.

Source: arXiv cs.AI | 2026-04-28

Loading related sources…