The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works
arXiv:2605.26246v1 Announce Type: new Abstract: Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on