GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
arXiv:2605.22812v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, exi