Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence
DGX agentarXiv:2605.19727v1 Announce Type: new Abstract: Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by com