ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation
DGX agentarXiv:2606.23104v1 Announce Type: new Abstract: On-policy distillation (OPD) improves LLM reasoning by training a student model on its own generated outputs, but standard OPD treats all student-genera