DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation
arXiv:2607.29078v1 Announce Type: new Abstract: On-policy distillation (OPD) trains student models on their own rollouts to reduce exposure bias. However, in multi-turn agent scenarios, early student