GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios
DGX agentarXiv:2606.06967v1 Announce Type: new Abstract: Generative policies provide expressive and multimodal action distributions, making them attractive for reinforcement learning (RL) in complex continuous