Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation
DGX agentarXiv:2606.00628v1 Announce Type: new Abstract: Self-distillation improves learning efficiency by rewriting reference answers as training data that better matches the model's own distribution. However