Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies
DGX agentarXiv:2607.05122v1 Announce Type: new Abstract: Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions an