Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
arXiv:2602.00807v2 Announce Type: replace Abstract: Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. H