Research
DINO_4D: Semantic-Aware 4D Reconstruction
arXiv:2604.09877v1 Announce Type: cross Abstract: In the intersection of computer vision and robotic perception, 4D reconstruction of dynamic scenes serve as the critical bridge connecting low-level g
arXiv:2604.09877v1 Announce Type: cross Abstract: In the intersection of computer vision and robotic perception, 4D reconstruction of dynamic scenes serve as the critical bridge connecting low-level geometric sensing with high-level semantic understanding. We present DINO_4D, introducing frozen DINOv3 features as structural priors, injecting semantic awareness into the reconstruction process to effectively suppress semantic drift during dynamic tracking. Experiments on the Point Odyssey and TUM-Dynamics benchmarks demonstrate that our method maintains the linear time complexity O(T) of its predecessors while significantly improving Tracking Accuracy (APD) and Reconstruction Completeness. DINO_4D establishes a new paradigm for constructing 4D World Models that possess both geometric precision and semantic understanding.
Related
- Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories
- RAM: Recover Any 3D Human Motion in-the-Wild
- 3D-Anchored Lookahead Planning for Persistent Robotic Scene Memory via World-Model-Based MCTS
Source: arXiv cs.AI | 2026-04-14