Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
DGX agentarXiv:2604.13016v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a core technique in the post-training of large language models, yet its training dynamics remain poorly unders