Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models
arXiv:2607.26513v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inhere