Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation
DGX agentarXiv:2606.03100v1 Announce Type: new Abstract: Recently, zero-shot 3D scene understanding via 2D Vision-Language Models (VLMs) has gained increasing research interest due to their promising spatial r