Applications

Is Knowledge Distillation Actually Greener? A Case Study in Machine Translation

arXiv:2602.09691v2 Announce Type: replace Abstract: Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student. In machine translation, KD is commonly evalua

DGX agentpaper
applicationsarxiv-cs-cl

arXiv:2602.09691v2 Announce Type: replace Abstract: Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student. In machine translation, KD is commonly evaluated through translation quality and inference efficiency, without jointly accounting for the environmental costs of producing and deploying the distilled system. We evaluate representative KD methods both on bespoke MT models and LLMs, by considering both translation quality and computational cost, using the Machine Learning Life Cycle Assessment tool, which accounts for costs throughout the KD model life cycle. Our key finding is that the deployment volume required to amortize KD is serving-dependent and can shift by several orders of magnitude under batching. We include actionable guidance for selecting, developing, and evaluating KD methods under quality and compute-induced constraints.

Related

Source: arXiv cs.CL | 2026-09-02

Loading related sources…