RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
DGX agentarXiv:2606.11709v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distri