SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
arXiv:2607.27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task