Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
DGX agentarXiv:2606.01621v1 Announce Type: new Abstract: Vision-language models (VLMs) have become a common foundation for vision-and-language navigation in continuous environments (VLN-CE). Yet most VLM-based