Safe Online Learning via Smooth Safety-Structured Policy Composition
DGX agentarXiv:2606.31320v1 Announce Type: new Abstract: Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics. Existing approaches t