Applications

V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

arXiv:2608.25308v1 Announce Type: new Abstract: Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understand

DGX agentpaper
applicationsarxiv-cs-cv

arXiv:2608.25308v1 Announce Type: new Abstract: Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understanding, and continuous action control. However, we reveal a critical limitation of VLA architectures: the action expert has limited access to the 3D geometric and 2D semantic information available in VLM features. This accessibility gap weakens perceptual grounding and limits performance on fine-grained robotic manipulation. To address this issue, we propose V-Link, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer. Specifically, V-Link learns complementary Spatial and Semantic Query representations within the VLM and injects them into Action DiT through asymmetric pathways. Semantic Queries complement the original VLM image tokens, whereas Spatial Queries provide dedicated geometric conditioning for spatially grounded action generation. Across LIBERO, LIBERO-Plus, and RoboTwin 2.0, our V-Link improves the average success rate over base model GR00T N1.6 by +1.9%, +31.2%, and +18.8%, respectively. On the AGIBOT A3 Ultra, V-Link further achieves gains of +20% and +24% on two real-world humanoid tasks.

Related

Source: arXiv cs.CV | 2026-08-27

Loading related sources…