Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
arXiv:2605.13632v1 Announce Type: cross Abstract: In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied