Safety
R3D: Revisiting 3D Policy Learning
arXiv:2604.15281v1 Announce Type: new Abstract: 3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe o
arXiv:2604.15281v1 Announce Type: new Abstract: 3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe overfitting, precluding the adoption of powerful 3D perception models. In this work, we systematically diagnose these failures, identifying the omission of 3D data augmentation and the adverse effects of Batch Normalization as primary causes. We propose a new architecture coupling a scalable transformer-based 3D encoder with a diffusion decoder, engineered specifically for stability at scale and designed to leverage large-scale pre-training. Our approach significantly outperforms state-of-the-art 3D baselines on challenging manipulation benchmarks, establishing a new and robust foundation for scalable 3D imitation learning. Project Page: https://r3d-policy.github.io/
Related
- LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
- HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
- EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
- Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers
- AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
- Action Images: End-to-End Policy Learning via Multiview Video Generation
Source: arXiv cs.CV | 2026-04-17