Model Releases
eta-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
arXiv:2607.28582v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reli
arXiv:2607.28582v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the eta=1 member of a broader policy-optimization family, where eta weights the KL penalty anchoring the student to a reference policy. This equivalence turns eta from an implicit value fixed at one into a controllable regularization parameter, yielding a more general formulation that trades off proximity to a reference policy against privileged teacher guidance. We introduce eta-OPSD and derive its optimal policy as a geometric interpolation between the reference policy and the privileged teacher. Directly optimizing this objective with reinforcement learning, however, would be costly and high-variance. Rather than optimize the RL objective directly, we turn its closed-form solution into a distillation target. Each value of eta selects a target along the reference-to-teacher path, which we implement efficiently by mixing their token-level logits. In this way, inexpensive distillation approximates the solution of expensive policy optimization. Return-to-go credit assignment further aligns token updates with the sequence-level objective while retaining the simplicity of OPSD. Experiments on mathematical reasoning benchmarks show that eta-OPSD consistently outperforms vanilla OPSD, improving optimization stability and downstream reasoning performance. Our results provide a principled route from self-distillation to policy optimization and back without sacrificing the efficiency that makes OPSD practical.
Related
- Predictable GRPO: A Closed-Form Model of Training Dynamics
- CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
- Self-Supervised On-Policy Distillation for Reasoning Language Models
- ReCo: Reweighting GRPO Against Distributional Concentration
- When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
Source: arXiv cs.LG | 2026-07-31