Safety

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning

arXiv:2511.18242v3 Announce Type: replace Abstract: Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal la

DGX agentpaper
safetyarxiv-cs-cv

arXiv:2511.18242v3 Announce Type: replace Abstract: Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) struggle with this setting, often generating plausible but visually inconsistent or weakly grounded responses. We introduce extbf{EgoVITA}, a framework that decomposes egocentric video reasoning into a structured extit{plan-then-verify} process. The model first generates an extbf{egocentric plan}: a causal sequence of anticipated actions from a first-person perspective. This plan is then evaluated by an extbf{exocentric verification} stage that uses third-person reasoning over the same video to verify its spatiotemporal and logical consistency, without exocentric video input. This decomposition enables cross-perspective feedback without requiring paired ego-exo supervision. To train this reasoning process, we adopt Group Relative Policy Optimization (GRPO) with two dense reward signals: one that grounds anticipated actions in subsequent visual observations and another that reinforces consistent third-person verification. extbf{EgoVITA} achieves state-of-the-art performance on egocentric reasoning benchmarks, outperforming Qwen2.5-VL-7B by mathbf{+7.7} on EgoBlind and mathbf{+4.4} on EgoOrient, while maintaining strong generalization on exocentric video tasks with only 52k training samples.

Source: arXiv cs.CV | 2026-07-01

Loading related sources…