Toward Sustainable Distributed LLM Inference: A Systems Synthesis and Research Agenda for an Energy-, Carbon-, and Cache-Aware llm-d Control Plane

arXiv:2609.05565 · cs.DC, cs.LG · Submitted 2026-09-03 · Read on arXiv

cs.DC, cs.LG

Submitted: 2026-09-03

Updated: 2026-09-03

Comments: 13 pages, 1 figure, 1 table. Systems synthesis and research agenda for sustainability-aware distributed LLM inference. No new experimental measurements are claimed; reported results are attributed to the cited work

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large language model (LLM) sustainability is increasingly a serving-systems problem, not only a training problem.

Terminology

Abstract

Large language model (LLM) sustainability is increasingly a serving-systems problem, not only a training problem. In production, energy and carbon impact depend on more than model size: workload shape, batching, key-value (KV) cache reuse, prefill/decode placement, model and accelerator choice, power state, geographic carbon intensity, and service-level objectives (SLOs) all matter. Recent systems papers study many of these factors separately. This paper connects those results and asks a practical engineering question: what do they imply when the decision point is a distributed inference control plane such as llm-d? The contribution here is synthesis, not a new set of benchmark results. Reported performance, energy, carbon, and cost improvements remain the results of the cited papers and systems. I group the literature into recurring design patterns and use those patterns to sketch a Sustainable Inference Control Plane (SICP) for llm-d. The proposed control plane would consider latency, energy, carbon, cache reuse, serving cost, and quality when routing and scaling, while keeping TTFT/TPOT SLOs as hard constraints. I also outline an evaluation framework based on SLO-satisfied goodput per joule and per gram CO2e, together with a reproducible experimental plan. The main observation from connecting the literature is that sustainable LLM inference is unlikely to come from one "green" model or one accelerator; it is more naturally treated as a control problem across model, phase, cache, hardware, replica, region, and time.

Related papers