Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
DGX agentarXiv:2607.28026v1 Announce Type: new Abstract: Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Se