Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs
arXiv:2606.31471v1 Announce Type: new Abstract: Open-vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision-language models, leaving the graph un