LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment
arXiv:2606.11221v1 Announce Type: new Abstract: We take a Gromov-Wasserstein perspective on Vision-Language-Action (VLA) learning, where the goal is to make the relational geometry of action represent