Model Releases
Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
arXiv:2606.13657v3 Announce Type: replace Abstract: On-policy distillation (OPD) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student-generate
arXiv:2606.13657v3 Announce Type: replace Abstract: On-policy distillation (OPD) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student-generated trajectories and dense token-level teacher supervision. Yet how this hybrid training regime shapes a model remains poorly understood. We characterize the sparsity and geometry of OPD parameter updates across several language and vision-language model pairs and application settings. OPD updates are small and coordinate-sparse at checkpoint precision, while remaining distributed across layers and modules. This sparse support is operationally meaningful: masked training on the discovered subnetwork nearly recovers full-training performance. At the matrix level, the updates are numerically full-rank but spectrally concentrated. Their visible supports avoid coordinates emphasized by the source's principal structure and favor low-magnitude source coordinates, while the source singular-value spectra change little. Together, these findings show that OPD exhibits important weight-space signatures of on-policy post-training despite using dense teacher supervision.
Related
- Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
- Geometric Self-Distillation for Reasoning Generalization
- Stable On-Policy Distillation through Adaptive Target Reformulation
- Trajectory-Refined Distillation
Source: arXiv cs.LG | 2026-07-31