Model Releases
Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance
arXiv:2606.04970v1 Announce Type: cross Abstract: We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding ext
arXiv:2606.04970v1 Announce Type: cross Abstract: We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding extit{when} to interrupt, and extit{how} to coach. However, progress is limited by the absence of large-scale, cross-domain benchmarks that reflect realistic conditions, particularly the common case in which users deviate from the expected step sequence. We address this gap with four contributions: extbf{(1)}we release extbf{EgoProactive}, a large-scale wearable-egocentric dataset for proactive procedural assistance with explicit Out-of-Plan (OOP) annotations and recovery steps; extbf{(2)}Pro, GPTwe augment five established benchmarks (Ego4D, EPIC-KITCHENS, EgoExo4D, HoloAssist, HowTo100M) into extbf{Proextsuperscript{2}Bench} under a unified proactive-guidance schema; extbf{(3)}3.1we propose a extbf{decoupled planner--interaction architecture} specialized for procedural state, visual cues, and recovery injection; extbf{(4)}4.6, Geminiwe introduce a post-training recipe that transfers across model families, validated by cross-backbone replication on Llama4 and Qwen-3.6-VL. In extensive experiments, our trained Llama-4 system substantially improves objective intervention quality over strong proprietary baselines (Claude Opus5.2) and open-weight baselines (Qwen3VL~235B) baselines across all six datasets. Oracle-plan experiments further show that, when plan quality is controlled, the trained duplex model produces high-quality guidance and large gains on Out-of-Plan recovery.
Source: arXiv cs.AI | 2026-06-04