Path-Space Mirror Descent for On-Policy Reinforcement Learning under the Generalized Schrodinger Bridge
DGX agentarXiv:2603.21621v2 Announce Type: replace Abstract: Classical on-policy algorithms such as PPO and mirror descent policy optimization provide stable proximal policy updates through tractable action li