Model Releases
CoRE: Weakly Supervised Coarse-to-Fine Risk Evidence Learning in Driving Videos
arXiv:2608.25344v1 Announce Type: new Abstract: Perceived risk in driving evolves over time and may be supported by specific scene entities, yet supervision is typically limited to coarse video-level
arXiv:2608.25344v1 Announce Type: new Abstract: Perceived risk in driving evolves over time and may be supported by specific scene entities, yet supervision is typically limited to coarse video-level judgments. Learning when supporting evidence emerges and which entities support a risk predictor would ordinarily require costly temporal- and entity-level annotations. We introduce extbf{CoRE}, a weakly supervised coarse-to-fine framework that learns fine-grained prediction support from coarse video supervision. CoRE first trains a video-level predictor and then freezes it. Structured interventions over candidate temporal regions or entity tracks measure how each candidate changes the coarse prediction, producing graded prediction-effect targets. These targets are distilled into a student that directly predicts temporal and entity support from the original video, without requiring interventions at inference. We evaluate this learning principle across three complementary settings: RISEE tests perceived-risk support from subjective clip-level judgments without temporal or entity-level risk annotations; DoTA provides independent temporal event annotations for evaluating weakly supervised traffic-anomaly localization; and UCF-Crime tests whether the same coarse-to-fine mechanism extends to a standard non-driving anomaly-detection benchmark. Across these settings, CoRE learns informative fine-grained support from coarse supervision, with strong temporal localization on DoTA and competitive performance on UCF-Crime. These results show that coarse video predictions can provide useful supervision for recovering the fine-grained evidence supporting them, without requiring corresponding fine-grained labels.
Related
- DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
- SUV: Future Scene Understanding as Video Generation for End-to-End Driving
- SceneMiner: Identity-Preserving Multi-Task Fine-Tuning for Unified BEV Scene Mining
- SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
- CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios
Source: arXiv cs.CV | 2026-08-27