Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
arXiv:2604.02546v2 Announce Type: replace Abstract: Pretraining 3D encoders by aligning with Contrastive Language Image Pretraining (CLIP) has emerged as a promising direction to learn generalizable r