Being-H0.7: A Latent World-Action Model from Egocentric Videos
arXiv:2605.00078v1 Announce Type: cross Abstract: Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to a