Research
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
arXiv:2602.14021v2 Announce Type: replace Abstract: Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motio
arXiv:2602.14021v2 Announce Type: replace Abstract: Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motion: static multi-view reconstruction systems assume a rigid world, whereas dynamic tracking frameworks rely on explicit ego-motion estimation or separate object motion models. In this work, we propose Flow4R, a unified framework that treats relative scene flow as the central representation linking 3D structure, camera ego-motion, and dynamic object motion. Given a two-view input, Flow4R employs a shared Vision Transformer to predict a compact, pixel-aligned property set comprising 3D point positions, scene flow, pose weights, and confidence maps. This flow-centric formulation allows local geometry and bidirectional motion to be jointly inferred in a single feedforward pass, eliminating the need for explicit pose regression heads or complex bundle adjustment. By training jointly on static and dynamic datasets, Flow4R achieves state-of-the-art performance on 4D reconstruction and tracking benchmarks, demonstrating the power of the flow-centric formulation for spatiotemporal scene understanding.
Related
- Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction
- SAMOFT: Robust Multi-Object Tracking via Region and Flow
- Learning Efficient 4D Gaussian Representations from Monocular Videos with Flow Splatting
- 4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
- No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos
Source: arXiv cs.CV | 2026-07-28