Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos
arXiv:2606.18955v2 Announce Type: replace-cross Abstract: Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotati