StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
DGX agentarXiv:2512.21970v2 Announce Type: replace Abstract: While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and show limited viewpoin