Learning Action Priors for Cross-embodiment Robot Manipulation
arXiv:2606.26095v1 Announce Type: cross Abstract: Most Vision-Language-Action (VLA) models build on a Vision-Language Model (VLM) backbone by attaching an action module and optimizing the full policy