Applications
Is Knowledge Distillation Actually Greener? A Case Study in Machine Translation
arXiv:2602.09691v2 Announce Type: replace Abstract: Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student. In machine translation, KD is commonly evalua
arXiv:2602.09691v2 Announce Type: replace Abstract: Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student. In machine translation, KD is commonly evaluated through translation quality and inference efficiency, without jointly accounting for the environmental costs of producing and deploying the distilled system. We evaluate representative KD methods both on bespoke MT models and LLMs, by considering both translation quality and computational cost, using the Machine Learning Life Cycle Assessment tool, which accounts for costs throughout the KD model life cycle. Our key finding is that the deployment volume required to amortize KD is serving-dependent and can shift by several orders of magnitude under batching. We include actionable guidance for selecting, developing, and evaluating KD methods under quality and compute-induced constraints.
Related
- Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
- Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging
- On Temperature-Constrained Non-Deterministic Machine Translation: Potential and Evaluation
- Knowledge Distillation for Low-Resource Open-source Text-to-SQL Model
Source: arXiv cs.CL | 2026-09-02