Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
DGX agentarXiv:2606.01247v1 Announce Type: new Abstract: Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has la