Research
CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation
arXiv:2511.16428v3 Announce Type: replace Abstract: Self-supervised surround-view depth estimation enables dense, low-cost 3D perception with a 360{eg} field of view from multiple minimally overlappin
arXiv:2511.16428v3 Announce Type: replace Abstract: Self-supervised surround-view depth estimation enables dense, low-cost 3D perception with a 360{eg} field of view from multiple minimally overlapping images. Yet, most existing methods suffer from depth estimates that are inconsistent across overlapping images. To address this limitation, we propose a novel geometry-guided method for calibrated, time-synchronized multi-camera rigs that predicts dense metric depth. Our approach targets two main sources of inconsistency: the limited receptive field in border regions of single-image depth estimation, and the difficulty of correspondence matching. We mitigate these two issues by extending the receptive field across views and restricting cross-view attention to a small neighborhood. To this end, we establish the neighborhood relationships between images by mapping the image-specific feature positions onto a shared cylinder. Based on the cylindrical positions, we apply an explicit spatial attention mechanism, with non-learned weighting, that aggregates features across images according to their distances on the cylinder. The modulated features are then decoded into a depth map for each view. Evaluated on the DDAD and nuScenes datasets, our method improves both cross-view depth consistency and overall depth accuracy compared with state-of-the-art approaches. Code is available at https://abualhanud.github.io/CylinderDepthPage.
Related
- Self-Improving 4D Perception via Self-Distillation
- Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
- 3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis
- One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
- A3-FPN: Asymptotic Content-Aware Pyramid Attention Network for Dense Visual Prediction
Source: arXiv cs.CV | 2026-04-14