Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
DGX agentarXiv:2605.13554v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations,