GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation
arXiv:2606.08530v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background s