The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference
cs.CL, cs.AI
Submitted: 2026-06-20
Updated: 2026-09-14
Comments: Accepted to EMNLP 2026 Main
Code: https://github.com/MichiganNLP/language-energy-dividehttps:
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large language models (LLMs) are increasingly deployed in multilingual settings, yet the energy costs of serving these models across different languages remain poorly understood.
Terminology
Abstract
Large language models (LLMs) are increasingly deployed in multilingual settings, yet the energy costs of serving these models across different languages remain poorly understood. We present a systematic study of inference energy consumption across languages with ML.Energy framework (Chung et al., 2026). We find striking disparities: energy consumption per output token varies by up to 8.3 times across languages, while total energy for a fixed set of requests varies by up to 179 times between the cheapest (English, 17.6 kJ) and the most expensive (Pashto, 3,147 kJ) languages. Our analysis shows that this disparity is driven by two compounding factors: (1) higher per-token energy costs for languages using complex or rare scripts, and (2) more tokens generated for low-resource languages. Moreover, we find a double cost + performance penalty: languages with the highest energy footprints also tend to achieve the lowest task accuracy. We reveal that the energy divide persists across models, hardware, and tasks, suggesting a systemic energy inequity in multilingual LLM deployment. Finally, we recommend that the community treat energy as a first-class evaluation axis, extend reporting checklists and model cards to include it, and adopt deployment-side mitigations for better energy efficiency.
Sources
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Where Do the Joules Go? Diagnosing Inference Energy Consumption
- Training Verifiers to Solve Math Word Problems
- No Language Left Behind: Scaling Human-Centered Machine Translation
- The Llama 3 Herd of Models
- Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
- The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
- Gemma 3 Technical Report
- ESTAR: Early-Stopping Token-Aware Reasoning For Efficient Inference
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering