VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
DGX agentarXiv:2505.20279v4 Announce Type: replace-cross Abstract: The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes,