Safety

RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI

arXiv:2602.07837v4 Announce Type: replace Abstract: Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-wo

DGX agentpaper
safetyarxiv-cs-ro

arXiv:2602.07837v4 Announce Type: replace Abstract: Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-world systems cannot be arbitrarily accelerated, cheaply reset, or massively replicated, suggesting that real-world policy learning is not merely an algorithmic problem, but inherently a systems problem. We present USER, a nderline{U}nified and extensible nderline{S}ystnderline{E}m for real-world online policy leanderline{R}ning. On the systems side, USER introduces a hardware abstraction layer for unified robot management and an adaptive communication plane that enables efficient cloud-edge training. On the learning side, USER adopts a fully asynchronous training framework, designs a persistent and cache-aware replay buffer, and provides extensible abstractions for rewards, algorithms, and policies. Experiments in both simulation and the real world demonstrate that USER supports multi-robot coordination, heterogeneous manipulators, cloud-edge training with large models, and long-running asynchronous training. Together, these capabilities establish USER as a unified and extensible systems foundation for real-world online policy learning.

Source: arXiv cs.RO | 2026-09-01

Loading related sources…