VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
DGX agentarXiv:2603.22003v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models typically map visual observations and linguistic instructions directly to robotic control signals. This 'black-b