Safety
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
arXiv:2604.23724v1 Announce Type: cross Abstract: Expressway video anomaly detection is essential for safety management. However, identifying anomalies across diverse scenes remains challenging, parti
arXiv:2604.23724v1 Announce Type: cross Abstract: Expressway video anomaly detection is essential for safety management. However, identifying anomalies across diverse scenes remains challenging, particularly for far-field targets exhibiting subtle abnormal vehicle motions. While Vision-Language Models (VLMs) demonstrate strong semantic reasoning capabilities, processing global frames causes attention dilution for these far-field objects and incurs prohibitive computational costs. To address these issues, we propose VIBES, an asynchronous collaborative framework utilizing VLMs guided by Bayesian inference. Specifically, to overcome poor generalization across varying expressway environments, we introduce an online Bayesian inference module. This module continuously evaluates vehicle trajectories to dynamically update the probabilistic boundaries of normal driving behaviors, serving as an asynchronous trigger to precisely localize anomalies in space and time. Instead of processing the continuous video stream, the VLM processes only the localized visual regions indicated by the trigger. This targeted visual input prevents attention dilution and enables accurate semantic reasoning. Extensive evaluations demonstrate that VIBES improves detection accuracy for far-field anomalies and reduces computational overhead, achieving high real-time efficiency and explainability while demonstrating generalization across diverse expressway conditions.
Related
- LLM-based Realistic Safety-Critical Driving Video Generation
- Neuromorphic Continual Learning for Sequential Deployment of Nuclear Plant Monitoring Systems
- DAIRE: A lightweight AI model for real-time detection of Controller Area Network attacks in the Internet of Vehicles
- Towards Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining
- Environmental Understanding Vision-Language Model for Embodied Agent
Source: arXiv cs.AI | 2026-04-28