Research
Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking
arXiv:2609.00924v1 Announce Type: cross Abstract: Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and spatial relationships. Multi-object trackers localiz
arXiv:2609.00924v1 Announce Type: cross Abstract: Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and spatial relationships. Multi-object trackers localize and associate objects primarily using appearance and geometry observed only in the image plane, inheriting these ambiguities. To address this limitation, we introduce PLANET, an end-to-end multi-object tracker designed to move beyond the image plane. As an enabling step, we lift existing 2D tracking datasets into 3D. We then form world-grounded queries by embedding reconstructed 3D scene geometry into the features and positional encodings used during query formation. An auxiliary 3D location prediction task further encourages the queries to encode object positions during training. A complementary dual-resolution temporal memory preserves this evidence across longer temporal gaps. As a result, PLANET achieves state-of-the-art performance across three diverse benchmarks.
Related
- DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
- FeedbackTrack: Visual-Cortex-Inspired Cross-Frame Feedback for Transformer Tracking
- DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
- World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video
Source: arXiv cs.AI | 2026-09-02