Is Knowledge Distillation Actually Greener? A Case Study in Machine Translation
cs.CL, cs.LG
Submitted: 2026-02-10
Updated: 2026-08-31
Comments: Accepted to EMNLP Main Conference 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student.
Terminology
Abstract
Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student. In machine translation, KD is commonly evaluated through translation quality and inference efficiency, without jointly accounting for the environmental costs of producing and deploying the distilled system. We evaluate representative KD methods both on bespoke MT models and LLMs, by considering both translation quality and computational cost, using the Machine Learning Life Cycle Assessment tool, which accounts for costs throughout the KD model life cycle. Our key finding is that the deployment volume required to amortize KD is serving-dependent and can shift by several orders of magnitude under batching. We include actionable guidance for selecting, developing, and evaluating KD methods under quality and compute-induced constraints.
Sources
- Fully Synthetic Data Improves Neural Machine Translation with Knowledge Distillation
- Carbontracker: Tracking and Predicting the Carbon Footprint of Training Deep Learning Models
- More than Carbon: Cradle-to-Grave environmental impacts of GenAI training on the Nvidia A100 GPU
- Distilling the Knowledge in a Neural Network
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering