Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems
DGX agentarXiv:2606.27797v1 Announce Type: cross Abstract: Knowledge Distillation (KD) enables training smaller student models under the guidance of larger teacher models, and the widely adopted TRL library im