InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

arXiv:2608.12915 · cs.DC, cs.AI · Submitted 2026-08-13 · Read on arXiv

Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos

University of Cyprus

cs.DC, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Author copy of paper published at 34th International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication System (MASCOTS2026)

Code: https://github.com/Azure/AzurePublicDataset

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: InFactPlanner is a trace-driven decision-support framework for what-if analysis of sustainable AI data center deployment for LLM inference across single and geo-distributed sites.

Terminology

Summary

InFactPlanner is a trace-driven decision-support framework for what-if analysis of sustainable AI data center deployment for LLM inference across single and geo-distributed sites. It combines query traces, hardware–model profiles, candidate site configurations, PUE/WUE parameters, renewable generation models, and time-varying grid carbon intensity to estimate power, energy, carbon emissions, water use, latency, and server utilization. The framework abstracts low-level serving effects into configurable hardware-model profiles, enabling rapid comparison of site selection, capacity placement, hardware, model, renewable integration, and routing choices. The authors validate the energy accounting pipeline by reproducing reference LLM inference energy estimates with less than 10% deviation, evaluate scalability across multiple data centers and server counts, and demonstrate scenario-driven decision analyses for hardware selection, renewable placement, geographic deployment, and carbon-aware routing. Results show that sustainability-optimal choices can differ from latency-optimal ones, and that the carbon value of deployment depends strongly on the local grid mix.

The paper's contributions are threefold: (1) an infrastructure-level sustainability analysis model for LLM inference that connects workload properties, hardware-model serving profiles, facility parameters, renewable generation, grid carbon intensity, and water use at a common time resolution; (2) implementation of this model in InFactPlanner, a configurable and modular trace-driven framework supporting single-site and geo-distributed what-if analysis across alternative siting, hardware, model, renewable, and routing decisions; (3) evaluation through accounting validation, scalability analysis, and decision-oriented case studies showing when latency-, energy-, and carbon-optimal deployment choices diverge.

The framework is organized as a trace-driven pipeline that transforms deployment descriptions and LLM inference workloads into sustainability and performance reports. It receives two main inputs: a deployment configuration encoded in YAML describing one or more data centers, and an input trace describing the expected LLM inference workload over time. The Trace Processing Stage prepares the workload, either using detailed per-query traces (with arrival time, input tokens, generated tokens) or generating per-query traces from aggregate rates using sampled or synthetic token distributions. The Simulation Engine assigns queries to data centers and servers according to the selected load-balancing policy (default round-robin, with custom strategies such as latency-aware or carbon-aware routing possible), then estimates query start times, completion times, active concurrency, and latency based on throughput, first-token latency, and concurrency limits. The Sustainability Modeling Stage enriches execution statistics with energy, carbon, and water estimates, aggregating server-level power demand at the data center level, scaling by PUE, subtracting local renewable generation from total demand to compute residual grid energy, aligning with location-specific carbon-intensity data for emissions, and estimating water consumption using WUE values. The Report Builder combines simulation and sustainability outputs into structured Pandas Dataframes.

For latency modeling, when constant TPS is True, query duration is estimated as duration = TTFT + (generated tokens / TPS). When constant TTFT is True, TTFT is taken directly from the hardware–model profile; otherwise it is computed dynamically as a function of query characteristics such as input tokens. When constant TPS is False, TPS varies over time according to concurrently active queries per second. The simulator maintains active queries and available concurrency slots per server, starting queries immediately if a slot is available or queuing them otherwise, capturing the effect of finite server capacity on waiting time and utilization.

For power accounting, instantaneous server power is modeled as Pnode(t) = xt(Pmax − Pidle) + Pidle, where xt is a sampled load-dependent factor from a truncated log-normal distribution parameterized using reported P5–P95 range. Query-level power is amortized across concurrent requests as Pquery(t) = Pnode(t)/n(t), where n(t) is the number of concurrent requests at time t. Facility-level server power applies PUE: Pserver(t) = Pnode(t) · PUE. Energy over an interval is computed as a discrete sum converted to watt-hours: ED = (1/3600) Σ Pserver(t). Per-query energy is estimated as Equery = (Duration/3600) · P̄query · PUE.

For renewable energy modeling, the framework focuses on solar energy with two options: math-based models that scale observed solar radiation according to configured PV capacity, and ML-based models using pre-trained regressors. Renewable generation time series are used to separate local supply from residual grid demand. Carbon emissions are estimated by multiplying time-step grid energy by time-aligned carbon intensity, where carbon intensity is computed as a weighted average of source-specific carbon factors: CI = Σ Ei fi / Σ Ei. Water use is computed as Water = (EDC/PUE) · WUEonsite + EDC · WUEoffsite.

Validation against real-system studies shows that InFactPlanner reproduces reference energy accounting estimates with deviations below 9% for DeepSeek-R1 on DGX H100 (median energy errors of 7.18% for traditional workloads and 8.71% for reasoning workloads), and median differences in energy and prompt duration of 1.9% and 4.5% respectively for DeepSeek-R1 on DGX B200. Scalability evaluation shows that for one-day traces with 1.24M to 9.92M queries, the framework completes in under 30 minutes across all configurations, with runtime increasing from 4 to 9 minutes when varying data centers from 1 to 8 (with 10 servers each), and from 4 to 28 minutes when varying servers per site from 10 to 80 (single data center). Memory usage was at most 25GB, showing it can run on commodity servers.

Scenario-based what-if case studies reveal: (1) For hardware–model pairings, Llama models consume up to 50% more energy and produce up to 50% higher carbon emissions than gpt-oss, while DGX B200 pairings increase environmental footprint by 35% due to higher system power but reduce mean latency by up to 77% and active queries per server by 63%. (2) For single-site RES integration, the estimated footprint is shaped by temporal alignment between workload demand, local RES generation, and residual grid dependence. (3) For multi-site deployments in Greece and Spain, placing RES in Greece reduces carbon emissions by 7.5% compared to Spain, because it displaces electricity from a more carbon-intensive grid, even though RES energy values differ by only 0.3% due to similar latitudes and timezones. (4) For carbon-aware routing, carbon-only routing achieves the largest carbon reduction (7% compared to round-robin) but increases latency by 22%, while balanced policy keeps latency within 2% of round-robin with only 1% carbon reduction, and carbon-dominant routing offers an intermediate trade-off (3% carbon reduction).

The paper acknowledges validity considerations: InFactPlanner targets deployment-level what-if analysis rather than low-level GPU serving simulation, abstracting fine-grained effects such as batching, KV-cache pressure, memory contention, and prefill/decode separation through configurable hardware–model profiles. Accuracy depends on input model fidelity, and results should be interpreted as comparative planning studies rather than exact production-footprint certification. The framework excludes embodied carbon and does not credit surplus RES exports.

Improvements for AI systems

Improvements to AI Systems:

  1. Carbon-Aware Inference Routing with Multi-Objective Optimization
  • Enhance AI serving systems to dynamically route queries across geo-distributed data centers using a tunable policy that balances carbon reduction, latency, and energy use. The improved system can adapt routing in real-time based on grid carbon intensity, renewable generation, and workload concurrency, achieving up to 7% carbon reduction with minimal latency penalty (e.g., <2% increase) by learning optimal trade-off weights from historical traces.
  1. Sustainability-Aware Hardware and Model Selection
  • Integrate a recommendation engine that, given a workload trace and deployment constraints, automatically selects hardware–model pairings (e.g., gpt-oss vs. Llama, DGX H100 vs. B200) to minimize carbon/water footprint while meeting latency Service Level Agreements (SLAs). The system can quantify trade-offs (e.g., 35% higher footprint for 77% lower latency) and suggest Pareto-optimal configurations.
  1. Renewable Integration Placement Optimizer
  • Add a planning module that, for multi-site deployments, computes the optimal placement of solar/wind capacity to maximize carbon displacement, considering local grid mix, timezone alignment, and workload temporal patterns. The improved system can recommend siting (e.g., Greece over Spain) that yields up to 7.5% additional carbon savings without significant energy yield differences.
  1. Real-Time Sustainability Forecasting and What-If Simulation
  • Enable AI systems to run continuous what-if simulations using live or forecasted traces, grid carbon intensity, and renewable generation. This allows operators to preemptively adjust routing, capacity, or renewable usage hours before carbon spikes, reducing operational emissions by 5–10% compared to reactive policies.
  1. Water-Aware Scheduling
  • Extend the routing and scheduling logic to include water consumption (WUE-based) as a constraint or objective, enabling AI systems to minimize water use in water-stressed regions while maintaining performance—useful for data centers in arid climates.
  1. Scalable Trace-Driven Capacity Planning
  • Use the framework’s scalability (handling 10M queries in <30 minutes) to build an AI-assisted capacity planner that tests thousands of deployment configurations (sites, servers, hardware) in parallel, identifying the most sustainable design under future workload growth and grid decarbonization scenarios.
  1. Uncertainty-Aware Decision Support
  • Incorporate the framework’s validation error margins (<10%) into AI-driven recommendations by outputting confidence intervals for carbon/energy estimates. The improved system can flag decisions where model uncertainty is high, prompting human review or fallback to conservative choices.
  1. Embodied Carbon and Lifecycle Extension
  • Extend the framework’s scope to include embodied carbon of hardware and renewable assets, allowing AI systems to optimize total lifecycle emissions (operational + embodied) for deployment choices—crucial for comparing frequent hardware refreshes vs. longer use.

Sources

Related papers