Safety

Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation

arXiv:2605.10118v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated exceptional general reasoning capabilities. However, their performance in embodied navigation remains hi

DGX agentpaper
safetyarxiv-cs-ro

arXiv:2605.10118v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated exceptional general reasoning capabilities. However, their performance in embodied navigation remains hindered by a scarcity of aligned open-world vision and robot control data. Despite simulators providing a cost-effective alternative for data collection, the inherent reliance on photorealistic simulations often limits the transferability of learned policies. To this end, we propose extit{extbf{S}andbox-extbf{A}bstracted extbf{G}rounded extbf{E}xperience} (extbf{extit{SAGE}}), a framework that enables agents to learn within a physics-grounded semantic abstraction rather than a photorealistic simulation, mimicking the human capacity for mental simulation where plans are rehearsed in simplified physics abstractions before execution. extit{SAGE} system operates via three synergistic phases: (1) extit{Genesis}: constructing diverse, physics-constrained semantic environments to bootstrap experience; (2) extit{Evolution}: distilling experiences through Reinforcement Learning (RL), utilizing a novel asymmetric adaptive clipping mechanism to stabilize updates; (3) extit{Navigation}: bridging the abstract policy to open-world control. We demonstrate that extit{SAGE} significantly improves planner-assisted embodied navigation, achieving a 53.21% LLM-Match Success Rate on A-EQA (+9.7% over baseline), while showing encouraging transfer to physical indoor robot deployment.

Source: arXiv cs.RO | 2026-05-12

Loading related sources…