SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
DGX agentarXiv:2605.11462v1 Announce Type: new Abstract: Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle