Safety
Self-Distilled RLVR
arXiv:2604.03128v2 Announce Type: replace Abstract: On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to pro
arXiv:2604.03128v2 Announce Type: replace Abstract: On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals for each sampled trajectory, in contrast to reinforcement learning with verifiable rewards (RLVR), which only obtains sparse signals from verifiable outcomes in the environment. Recently, the community has explored on-policy self-distillation (OPSD), where the same model serves as both teacher and student, with the teacher receiving additional privileged information such as reference answers to enable self-evolution. This paper demonstrates that learning signals solely derived from the privileged teacher result in severe information leakage and unstable long-term training. Accordingly, we identify the optimal niche for self-distillation and propose extbf{RLSD} (extbf{RL}VR with extbf{S}elf-extbf{D}istillation). Specifically, we leverage self-distillation to obtain token-level policy differences for determining fine-grained update magnitudes, while continuing to use RLVR to derive reliable update directions from environmental feedback (e.g., response correctness). This enables RLSD to simultaneously harness the strengths of both RLVR and OPSD, achieving a higher convergence ceiling and superior training stability.
Related
- Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
- GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA
- Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs
- Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
- FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
Source: arXiv cs.LG | 2026-04-10