Safety
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
arXiv:2602.00181v3 Announce Type: replace-cross Abstract: Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat t
arXiv:2602.00181v3 Announce Type: replace-cross Abstract: Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat this task as a black-box classification, often confusing physically distinct motions by relying on superficial visual patterns rather than geometric cues. We present extbf{CamReasoner}, a framework that reformulates camera movement understanding as a structured inference process to bridge the gap between perception and cinematic logic. Our approach centers on the Observation-Thinking-Answer (O-T-A) paradigm, which compels the model to articulate spatio-temporal observations and reason about motion patterns within an explicit reasoning block. To instill this capability, we construct a Large-scale Inference Trajectory Suite comprising 18k SFT reasoning chains and 38k RL feedback samples. To the best of our knowledge, extbf{we are the first to employ RL for logical alignment in camera movement understanding}, ensuring motion inferences are grounded in structured visual reasoning rather than contextual guesswork. Built upon Qwen2.5-VL-7B, CamReasoner-7B improves binary classification accuracy from 73.8% to 78.4% and VQA accuracy from 60.9% to 74.5% over its backbone, consistently outperforming both proprietary and open-source baselines across multiple benchmarks.
Related
- Visually-Guided Policy Optimization for Multimodal Reasoning
- MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models
- MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
- Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization
- Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine
Source: arXiv cs.AI | 2026-04-15