Aging-Aware Online Distributed Scheduling for Lifecycle Carbon Reduction in Geo-Distributed Data Centers

arXiv:2610.00605 · eess.SY, cs.SY · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Aging-Aware Online Distributed Scheduling for Lifecycle Carbon Reduction in Geo-Distributed Data Centers".

Rosa: The rapid proliferation of data centers has led to massive energy demand and carbon emissions, posing significant sustainability challenges.

Dev: First, who's behind it and why it matters.

Paper summary: Dev: To wrap up on the paper "Aging-Aware Online Distributed Scheduling for Lifecycle Carbon Reduction in Geo-Distributed Data Centers," the authors successfully presented a comprehensive modeling framework that integrates electricity flow, workload flow, and degradation flow to evaluate long-term carbon costs from server degradation.

Rosa: They achieved a reduction of thirteen point zero percent in total carbon emissions and twelve point six percent lower operational costs when compared against the benchmarks they tested. This result shows that combining workload scheduling with an utilization-dependent exponential aging model can deliver tangible financial and environmental benefits.

Taro: The authors are essentially proving that by accounting for how hardware degrades under high utilization, you can proactively adjust your scheduling to minimize future embodied carbon costs. This moves the focus from just today's energy use to the entire lifespan of the data center infrastructure.

Dev: It’s clear that this work is about creating a unified optimization problem that minimizes both immediate financial costs and long-term lifecycle carbon emissions through an online scheduling approach. The methodology relies on transforming the complex inter-temporal constraints into a sequence of per-slot deterministic problems using an online Lyapunov optimization framework.

Rosa: In simple terms, the paper shows that you can achieve significant savings by making your distributed data centers smarter about how they schedule workloads based on both current energy prices and the predicted physical wear of their servers over time.

Taro: The real impact is suggesting that for autonomous systems operating in these environments, making decisions that consider hardware lifespan and carbon cost simultaneously becomes a necessary component of intelligent operation.

Dev: We should keep an eye on how they validate this framework in real-world scenarios; the success hinges on whether the online adaptation mechanisms, like the Time-Varying Queue Shifting, hold up under unpredictable real-world fluctuations.

Conclusion: Rosa: So, we've seen how they model the carbon lifecycle using aging—what do you make of that title?

Dev: I think "Aging-Aware Online Distributed Scheduling" tells us immediately that this isn't just a static optimization problem; it’s dealing with dynamic physical reality and requiring real-time adjustments.

Taro: From an autonomy standpoint, the "Online" part is crucial because it implies the system has to react when things go wrong or change unexpectedly in the environment.

Rosa: Exactly, and I'm wondering if this model can actually function outside of a controlled lab setting for a long duration?

Dev: That’s a big question; we need to know how robust these Lyapunov functions are when faced with unexpected hardware failures or drastic shifts in grid pricing.

Taro: The real world is messy, Rosa; I'm curious what happens when the predicted degradation model deviates significantly from reality under severe stress.

Rosa: It seems the authors are arguing that by incorporating this wear into the decision-making process, we can get a much better long-term picture of our environmental footprint for data centers.

Dev: And if they’re managing to cut those operational costs by twelve point six percent while reducing emissions by thirteen percent, that suggests a very tight balance between efficiency and longevity.

Taro: It points toward a future where resource allocation in distributed systems isn't just about minimizing today's power draw but about designing infrastructure for its entire useful life.

Rosa: If this concept scales up to massive, geographically dispersed data centers, the potential for reducing the overall carbon load across industries is pretty substantial.

Dev: We should focus on the latency implications; if these online adjustments introduce too much control loop overhead, it defeats the purpose of real-time optimization.

Taro: That’s a valid concern for any autonomous system; we need to ensure that optimizing for lifecycle cost doesn't sacrifice immediate performance metrics.

Junyu Lin, Wenjie Liu, Shunbo Lei

Guangzhou University · The Chinese University of Hong Kong · UC Berkeley

eess.SY, cs.SY

Submitted: 2026-09-30

Updated: 2026-09-30

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 78/100

The gist: The rapid proliferation of data centers has led to massive energy demand and carbon emissions, posing significant sustainability challenges.

Key concepts

Electricity Flow Modeling
This models the cost of buying electricity as a function of how much power is drawn from the grid, the current price per unit, and the time slot. It helps determine the immediate financial impact of energy consumption on data center operations.
Degradation-Informed Embodied Carbon Cost Model
This links how hard servers are worked (utilization) to their physical aging using an exponential model. This aging is then translated into a monetary cost based on embodied carbon, accounting for the environmental impact of hardware wear over time.
Lyapunov Optimization Framework
This mathematical technique transforms a complex long-term optimization problem into a sequence of simpler, real-time problems. It uses queue variables to manage constraints across different time periods, allowing for online, adaptive decision-making in distributed systems.

Terminology

Summary

The rapid proliferation of data centers has led to massive energy demand and carbon emissions, posing significant sustainability challenges. The proposed framework addresses these limitations by combining workload scheduling with an utilization-dependent exponential aging model to evaluate long-term carbon costs from server degradation, achieving up to 13.0% lower carbon emissions and 12.6% lower operational costs compared with benchmarks.

The gist

The paper proposes a comprehensive carbon life-cycle modeling framework for distributed data centers that jointly accounts for operational carbon emissions and hardware-aging-induced embodied carbon emissions, resulting in a reduction of 13.0% in total carbon emissions and 12.6% in operational costs relative to benchmarks.

How it works

The proposed solution is structured around a lifecycle model that integrates three coupled flows: electricity flow, workload flow, and degradation flow. This framework aims to jointly minimize economic costs and lifecycle carbon emissions by optimizing the control variables across geographically distributed data centers.

  1. Electricity Flow Modeling: This involves modeling the electricity purchase cost from the grid, which is defined as a function of active power drawn from the grid, electricity price, and duration of each time slot:

f g i,t = λ i,t Pg i,t ∆t.

  1. Degradation-Informed Embodied Carbon Cost Model: This model links workload allocation to physical server aging and embodied carbon cost. It defines the resource utilization level ui,k,t as the ratio of assigned workload to total available processing capacity:

ui k,t ≜ Wi k,t / (N A i,k S rate i)

The load-dependent physical wear is mapped using an exponential aging model:

Luse i,k,t+1 = Luse i,k,t + ηi,k exp(γi k ui k,t ∆t) t ∈ [z]

The instantaneous embodied carbon expenditure for a cluster is calculated as:

Eemb i k,t = C M i k∆li k,t (8), where the fractional lifetime consumption is calculated as:

∆li k,t = ηi k exp(γi kui k,t)∆t / Llife i k.

The total embodied carbon cost for data center i at time slot t is aggregated across all active servers and converted into a monetary cost via the carbon price αC:

f c,m i,t = αC X k∈Ki N A i,k,t Eemb i k,t (9).

How it works (Continued)

The total carbon-related financial objective for data center i at time slot t is established as the sum of operational and embodied carbon costs:

f c i,t = f c o i,t + f c m i,t (10).

How it works (Continued)

The long-term optimization problem P1 is transformed into a sequence of per-slot deterministic optimization problems P2 using an online Lyapunov optimization framework to decouple the inter-temporal constraints. This involves defining three queue-based state variables: Batch Workload Queue (S D), Thermal Virtual Queue (τ˜i,t), and Battery Virtual Queue (E˜i,t).

The Time-Varying Queue Shifting (TVQS) mechanism adjusts the battery virtual-queue shift parameter ε i,t based on real-time electricity prices and renewable generation:

ε i,t = -Emax b i θ i,t (35), where the operational incentive factor θ i,t is defined by:

θ i,t = min (1, βθ / (P r i t P r rated i + P r i t + λ max i − λi,tλ max i !))

The unified Lyapunov function L(Θt) is defined as:

L(Θt) = 1/2 X i∈I [α i(S D i,t) squared + δ iE˜ squared i,t + ν iτ˜ squared i,t] (36).

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to AI systems, categorized by the capability they will gain:


) Carbon-Aware, Lifecycle-Optimized Resource Orchestration System (CLOROS): This system will move beyond traditional energy efficiency by integrating hardware degradation directly into its scheduling logic.

The improved system will dynamically allocate workloads across geographically distributed data centers not just based on real-time operational costs or immediate energy prices, but also by predicting the long-term embodied carbon cost associated with accelerated hardware aging (utilization-dependent thermal stress).

) Predictive, Adaptive Battery Management System: This system will replace static battery scheduling with a mechanism that actively anticipates future electricity price and renewable generation fluctuations.

The improved system will utilize the Time-Varying Queue Shifting (TVQS) mechanism to proactively charge or discharge energy storage based on forecasted electricity prices and renewable availability, leading to significantly more optimized energy arbitrage and reduced operational costs.

) Privacy-Preserving Distributed Workload Migration Network: This system will enable high-efficiency workload balancing across multiple independent data centers without compromising the security or privacy of local user/application data.

The improved system will employ the Zero-Sum Perturbation ADMM (ZSP-ADMM) framework to coordinate workload migration decisions across DCs in a way that masks individual migration volumes, ensuring global load balance is maintained while preventing adversaries from inferring sensitive local workload profiles.

) Hardware Health-Aware Cluster Selection Engine: This system will optimize server placement based on the predicted lifespan of different hardware components.

The improved system will prioritize allocating workloads to server clusters within a data center that exhibit lower degradation rates (as quantified by utilization and aging factors), thereby extending the effective operational life of the physical hardware and drastically reducing future embodied carbon emissions from replacement.

) Joint Economic and Environmental Optimization Controller: This is a holistic controller designed to minimize both immediate financial expenditure and total lifecycle carbon footprint simultaneously.

The improved system will solve a single, unified optimization problem (P1/P3) that balances the instantaneous operational costs (energy, transfer fees, operational carbon) against the long-term embodied carbon costs derived from hardware wear and tear, resulting in a globally optimal low-carbon deployment strategy.

This improved AI system can perform the following specific actions:

  1. Identify and migrate latency-critical interactive workloads (IWs) to data centers offering the lowest combination of immediate electricity cost and future hardware degradation risk.

  2. Manage battery energy storage in real-time, charging during periods of low carbon intensity or low electricity prices, and discharging strategically during high-cost/low-renewable peak hours, adapting its strategy dynamically via TVQS.

  3. Balance the total workload across a global network of data centers by coordinating migration decisions using ZSP-ADMM, ensuring that no single DC's local workload profile is exposed to external coordination mechanisms.

  4. Select the optimal cluster within a data center for any given task based on its utilization level and inherent hardware aging coefficient, ensuring that high-utilization tasks are directed toward younger hardware to minimize accelerated degradation and subsequent embodied carbon liability.

Abstract

The rapid proliferation of data centers (DCs), driven by cloud computing and artificial intelligence (AI), has led to massive energy demand and carbon emissions, posing significant sustainability challenges. Carbon-aware optimization in geographically distributed data centers has been widely studied. Most existing approaches mainly focus on operational carbon emissions from server usage. However, existing literature often ignores workload-induced thermal stress, which accelerates nonlinear hardware degradation. This leads to more frequent server replacements and ultimately increases embodied carbon emissions. To address these limitations, we propose a comprehensive carbon life-cycle modeling framework for distributed data centers. Apart from operational carbon emissions, this work combines workload scheduling with a utilization-dependent exponential aging model to evaluate long-term carbon costs from server degradation. In order to solve the proposed optimization model in an online and privacy-preserving manner, an enhanced Lyapunov framework with time-varying queue shifting (TVQS) is first introduced to handle system uncertainties. Then, a zero-sum perturbation-based alternating direction method of multipliers (ZSP-ADMM) framework is developed to enable distributed coordination across geographically separated data centers while protecting locally exchanged workload information. Simulation results demonstrate that the proposed approach achieves up to 13.0% lower carbon emissions and 12.6% lower operational costs compared with benchmarks.

Sources

Related papers