Antifragile perimeter control: Thriving on disruptions through reinforcement learning

arXiv:2402.12665 · eess.SY, cs.SY · Submitted 2024-02-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Antifragile perimeter control".

Rosa: The optimal operation of transportation systems is often susceptible to unexpected disruptions, and many established control strategies reliant on mathematical models can struggle with real-world disruptions,

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So we're diving into "Antifragile perimeter control: Thriving on disruptions through reinforcement learning," and I'm really curious what that title implies about the approach they took. It sounds like they are moving away from just trying to keep things stable and aiming for something more dynamic when things get rough.

Dev: Yeah, it definitely suggests a system designed not just to withstand shocks but actually benefit from them, which is a significant shift in control philosophy compared to standard robustness studies. The authors are Linghang Suna, Michail A. Makridisa, Alexander Gensera, Cristian Axenieb, Margherita Grossic, and Anastasios Kouvelasa from ETH Zurich and the Technische Hochschule Nürnberg in Germany.

Taro: I'm interested in how they framed this concept of antifragility versus the terms like resilience or reliability that we usually see in risk engineering literature. It sounds like they are proposing a specific mechanism for how a system should respond to adversarial events.

Rosa: Exactly, and given the background we have on other papers, I wonder if this paper offers a concrete framework for how learning algorithms can embody that philosophy in real-time traffic management scenarios.

Dev: The core idea is integrating antifragility directly into the learning strategy to optimize urban road network operations specifically when disruptions occur. It’s not just about surviving the bad state; it’s about enhancing performance during those stressful periods, which is what they're aiming for with this paper.

Taro: That's where I get excited because when we think about autonomous systems in unpredictable environments, we need agents that don't just maintain a baseline function but actively improve their operational capability when the environment degrades.

The paper's summary: Rosa: So, summarizing what this paper actually proposes for the "Antifragile perimeter control: Thriving on disruptions through reinforcement learning," it seems they are using deep reinforcement learning to manage traffic flow across cordon-shaped networks, but with a unique twist involving antifragility principles.

Dev: Right, they are incorporating modules composed of traffic state derivatives and redundancy directly into the deep reinforcement learning algorithm to achieve this enhancement under disruptions. They’re not just using standard RL; they're building something designed to thrive when the system is stressed.

Taro: I see that they are explicitly designing this mechanism through state representation augmentation and reward function shaping, which suggests a very deliberate design choice rather than an emergent property of the learning process alone.

Rosa: That deliberate design is what makes it interesting; they’re tackling both fragile performance issues and observability problems simultaneously by building in these antifragile components.

Dev: Precisely, and they are testing this on a cordon-shaped transportation network and even a real-world case study with actual data to see if the theory translates into practical gains for traffic control.

Taro: That evaluation is crucial because it shows whether this theoretical framework holds up when we move from idealized simulations to messy, real-world data constraints.

The paper's improvements: Rosa: If we look at the specific technical enhancements mentioned in "Antifragile perimeter control: Thriving on disruptions through reinforcement learning," the main improvements are centered around how they handle state information and reward signals within the RL framework.

Dev: They introduce augmenting the state space with first- and second-order derivatives of traffic states, like n ij,k and squared n ij,k, which gives the agent predictive power about congestion rates.

Taro: That derivative information is key because it allows the AI to anticipate changes in flow—the rate of change—which should let it make anticipatory control adjustments instead of just reacting to what's happening right now.

Rosa: And they pair that with a sophisticated reward function featuring a damping term, r dam,k, which penalizes rapid oscillations in control actions, and a redundancy term, r red,k.

Dev: That redundancy term builds up system redundancy specifically to ensure the agent isn't overly dependent on one perfect set of conditions; it’s designed to make the system thrive by preparing it for unexpected variations.

Taro: The damping term is interesting because it directly addresses stability concerns in physical systems, which is a major consideration when deploying control strategies in traffic infrastructure.

Conclusion: Rosa: So, wrapping up the findings from "Antifragile perimeter control: Thriving on disruptions through reinforcement learning," the main conclusion is that this proposed algorithm outperforms baselines significantly under increasing demand and supply disruptions.

Dev: They found performance gains reaching up to twenty-seven point six percent and even forty-one point nine percent over baseline RL algorithms when faced with maximal demand and supply shocks, which is quite substantial for a control system under stress.

Taro: I noticed the evaluation uses distribution skewness as a quantitative indicator of antifragility, showing that their ultimate skewness at zero point four three is much better than the baselines' values like zero point eight four or zero point eight six, indicating relative antifragility against them in many scenarios.

Rosa: And they also showed it’s effective under limited observability using real-world data constraints, achieving a performance gain of about four point eight percent on average and an average skewness of zero point three nine in that scenario.

Dev: That trade-off under limited observability is important because it shows they can still extract value even when not having perfect information, which is a practical reality for deployed systems to consider.

Taro: I think the implication here is that this concept has broad applicability beyond just traffic engineering, suggesting it could be useful in any control system exposed to disruptions across different disciplines.

Institute for Transport Planning and Systems, ETH Zurich, Zurich, Switzerland · Computer Science Department and Center for Artificial Intelligence, Technische Hochschule Nürnberg, Nürnberg, Germany · Intelligent Cloud Technologies Lab, Huawei Munich Research Center, Munich, Germany

eess.SY, cs.SY

Submitted: 2024-02-20

Updated: 2026-09-24

Comments: 36 pages, 17 figures

Journal ref: Reliability Engineering & System Safety. 277 (2027) 113194

DOI: 10.1016/j.ress.2026.113194

Code: https://github.com/DongqinZhou/C-D-RL

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: The optimal operation of transportation systems is often susceptible to unexpected disruptions, and many established control strategies reliant on mathematical models can struggle with real-world

Key concepts

Antifragile Perimeter Control
This approach moves beyond simply surviving shocks by designing a system that actively benefits from disruptions. It aims to enhance operational capability when the environment degrades, rather than just maintaining stability.
Deep Reinforcement Learning (DRL)
The paper uses DRL to manage traffic flow across cordon-shaped networks. They enhance standard RL by incorporating modules with traffic state derivatives and redundancy directly into the learning algorithm.
State Space Augmentation
The authors augment the state space by including first- and second-order derivatives of traffic states. This provides the AI with predictive power regarding congestion rates, allowing for anticipatory control adjustments.
Reward Function Shaping
The reward function includes a damping term to penalize rapid oscillations in control actions and a redundancy term to build system redundancy. This design ensures the agent thrives by preparing for unexpected variations.

Terminology

Summary

The optimal operation of transportation systems is often susceptible to unexpected disruptions, and many established control strategies reliant on mathematical models can struggle with real-world disruptions, leading to significant divergence from their anticipated efficiency. This study integrates the cutting-edge concept of antifragility with learning-based traffic control strategies to optimize urban road network operations under disruptions. Antifragile systems not only withstand and recover from stressors but also thrive and enhance performance in the presence of such adversarial events. Incorporating antifragile modules composed of traffic state derivatives and redundancy, a deep reinforcement learning algorithm is developed. Subsequently, it is evaluated in a cordon-shaped transportation network and a case study with real-world data. Promising results highlight that the proposed algorithm provides: (i) superior performance achieving up to 27.6% and 41.9% performance gain over baselines under increasing demand and supply disruptions, (ii) lower distribution skewness under disruptions, demonstrating its relative antifragility against baselines, (iii) effectiveness under limited observability due to real-world data availability constraints, and (iv) the robustness and transferability to be combined with various state-of-the-art RL frameworks. The proposed antifragile methodology is generalizable and holds potential for applications beyond traffic engineering, offering integration into control systems exposed to disruptions across various disciplines.

The key contributions of this research can be summarized as follows:

  1. Conceptual contribution: We distinguish the concept of antifragility as opposed to other related and commonly used terms in the risk engineering and transportation domains.

  2. Methodological contribution: We introduce how antifragility can be incorporated into RL algorithms to achieve superior performance compared to benchmark methods, tackling both fragile performance and observability issues.

  3. Evaluatory contribution: We adopt a skewness-based quantitative indicator to showcase the antifragile properties of our proposed algorithm under increasing disruptions.

  4. General applicability: We further validate the effectiveness of the proposed antifragile module on other state-of-the-art RL algorithms and with a real-world case study.

The paper distinguishes the concept of antifragility from related terms such as robustness, resilience, reliability, stability, and adaptiveness. Robustness is concerned with assessing a system’s capacity to preserve its initial state and resist performance deterioration under minor disturbances, while resilience emphasizes the ability and speed of a system to recover from major disruptions to the original state. Reliability has various meanings, such as travel time reliability or an assessment indicator for performance loss. Adaptiveness is defined as the capacity of a system to modify its own characteristics to maintain autonomous function across diverse environments. Antifragility, introduced by Taleb (2012), emphasizes the concave response of the system under increasing disruptions, which can be mathematically formulated with Jensen’s inequality E[g(X)] ≤ g(E[X]). The authors seek to develop antifragile solutions that mitigate such fragility through perimeter control.

The problem formulation studies perimeter control between two homogeneous cordon-shaped urban networks. The total number of vehicles in region i at time t is denoted as ni (t), while the Origin-Destination (OD) from region i to region j can be further divided as nij (t). The inner and outer regions are assumed to have different MFDs to represent the capacity difference of accommodating vehicles within the network between the city center and the surrounding region, which are defined as Gi (ni (t)). The total trip completion rate Mi (t) for region i at time t can be determined through the corresponding MFD, which comprises both intraregional trip completion Mii (t) and interregional transfer flow Mij (t). The percentage of the transfer flow Mij (t) allowed to pass across the boundary is regulated by the perimeter controllers, denoted as uij (t) with i, j ∈ [1, 2].

The objective function for control-based strategies is J = max Mii(t)dt / (3). The reward function for the proposed antifragile RL-based algorithms Jaf is described in discrete form as:

Jaf = max uij(k) / (rcom,k + rdam,k + rred,k)

The state space S for the RL algorithm is defined differently depending on observability:

Idealized full observability: sk = [nij,k, ∆nij,k, ∆2 nij,k, Mij,k]

Real-world limited observability: sk = [ni,k, ∆ni,k, ∆2 ni,k, Mij,k]

The reward function incorporates two additional terms:

  1. The damping term rdam,k is introduced to penalize potential oscillatory actions, defined as:

rdam,k = −ξ1 uij,k − uij,k−1 ξ2, i ̸= j and k ̸= 1

  1. The redundancy term rred,k functions as an additional term to build up redundancy in the system, summarized as:

rred,k = Hk +∆Hk in discrete form.

The first difference Hk is defined as:

Hk = Σ i=1,2 ∆Hi,k = Σ i=1,2 Hi,k = ωh · hi,k · αi,k · f (ni,k, ni crit, ni cap)

The second difference ∆Hk is defined as:

∆Hi k = ω∆h · ∆hi k · f (ni,k, ni crit, ni cap)

The performance evaluation uses the Total Time Spent (TTS), calculated by adding up the number of vehicles within the network at each second of the simulation. To quantify antifragility, we calculate the distribution skewness based on the TTS sampled from the last Nepisode = 25 incremental episodes, where a negative skewness indicates a longer or fatter left tail of the distribution and thus a higher degree of concavity in the performance function, which showcases antifragility.

Results show that under idealized full observability, "the proposed antifragile RL algorithm exhibits both superior performance and reduced variance, indicated by its significantly narrower shaded area. Furthermore, the performance curve of the proposed method also appears less convex than the other algorithms. The average performance gain is 10.7% and reaches 27.6% at episode 75 under incremental demand disruption. The distribution skewness for the antifragile RL algorithm achieves an ultimate skewness at 0.43," compared to a baseline RL skewness of 0.84 and MPC's value of 0.86, indicating its relative antifragility against baselines.

Under supply disruptions, the proposed algorithm shows noticeably lower skewness across episodes 50 − 75 compared to the other approaches, with a final skewness of 0.54, whereas the skewness of the other three methods approaches or exceeds 1.0. This indicates that the proposed algorithm is only applicable under typical disruptions, while not suitable for handling catastrophic scenarios where the network rapidly descends into gridlock.

Under real-world limited observability, the performance curve of the antifragile RL algorithm under limited observability in light teal falls between that of the baseline RL and the antifragile RL with full observability, indicating a performance trade-off. The TTS difference between the two algorithms under limited observability yields an average performance gain of 4.8%. The average skewness for this scenario is 0.39, which lies between the skewness observed with full observability (0.24) and that of the baseline RL under limited observability (0.41)."

In conclusion, "this work is the first of its kind to pioneer the application of antifragility in the operation and control of transportation systems, continuously improving the urban network performance under unforeseen disruptions with learning-based algorithms. The study validates that the proposed algorithm exhibits greater antifragile properties with the lowest skewness among all the methods examined and delivers increasingly better performance as disruptions intensify. Results demonstrate the broad applicability of our proposed approach to other contemporary baselines, including TD3 and SAC. Finally, the work is validated with a real-world case study based on realistic demand, supply, and disruptions from the city center of Zurich. The authors acknowledge that real-world operation with the proposed antifragile RL-based algorithm can be more accurately validated through traffic microsimulation software like SUMO."

The final quantitative results show that the proposed algorithm achieves 27.6% and 41.9% performance improvements over the baseline RL algorithm under maximal demand and supply disruptions. The performance gain under supply disruption is shown to reach a ultimate performance improvement of 41.9%. Furthermore, in real-world limited observability, the trade-off is quantified as an average of 4.6%. Finally, the final skewness achieved by the proposed algorithm under supply disruptions is 0.46 compared to 1.04 from MPC, 0.92 from MPC-MHE, and 0.96 from the baseline RL algorithm. The performance and antifragile properties are studied under real-world limited observability, demonstrating that a certain degree of performance can be traded for using only accessible sensor measurements in reality. The study concludes that the concept is generic enough to be extended to other traffic control systems and potentially to control systems in general that are subject to growing disruptions. (Page 30)


Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed this paper, Antifragile perimeter control: Anticipating and gaining from disruptions with reinforcement learning. The core contribution is the development of an RL-based traffic control algorithm that incorporates concepts of antifragility—specifically through state representation augmentation (derivatives) and reward function shaping (damping and redundancy terms)—to enhance performance during disruptions.

Here are the specific improvements to AI systems derived from this research, along with what the improved AI system can achieve:


)1. State Representation Augmentation for Derivative Information

The paper introduces incorporating first-order and second-order derivatives of traffic states (vehicle accumulation) into the RL state space, moving beyond standard state vectors like vehicle counts or OD matrices.

"we propose a state representation incorporating both the first- and second-order derivatives of vehicle accumulation, which can be computed as the first and second differences from the vehicle accumulation [∆nij,k, ∆2 nij,k] between two consecutive time steps in the discrete form." (Section 4.3)

]

The improved AI system can achieve:

  • Anticipatory Control: The agent will not only react to current traffic congestion but will anticipate the rate of change (first derivative, velocity of congestion) and the acceleration/deceleration (second derivative, curvature of the MFD). This allows for proactive adjustments to perimeter control variables before a critical threshold is crossed.

  • Mitigation of Oscillations: By explicitly modeling these derivatives in the state space, the agent gains richer information to better stabilize its actions, leading to smoother control outputs (as validated by the damping term).

)2. Antifragile Reward Shaping for Proactive Adaptation

The paper introduces a novel reward function that includes specific terms designed to promote antifragility: a damping term and a redundancy term.

the reward is defined with additional two terms, i.e., the damping term rdam,k and the redundancy term rred,k (Section 7)

rdam,k = −ξ1 uij,k − uij,k−1 ξ2, i ̸= j and k ̸= 1 (Eq. 10 - Damping term)

rred,k = Hk +∆Hk in discrete form (Eq. 11 - Redundancy term)

]

The improved AI system can achieve:

  • Antifragile Response: The redundancy term actively builds system redundancy, ensuring the agent is not overly optimized for a single critical state but is instead prepared to handle unexpected variations. This makes the system thrive (gain performance) when facing stressors rather than merely surviving them.

  • Stability Under Uncertainty: The damping term penalizes excessive or rapid changes in control actions, preventing unstable oscillations in physical systems (like traffic signals), leading to more reliable and practical deployment, especially when operating under real-world limited observability.

)3. Robustness Against Observability Constraints (Limited Data)

The methodology explicitly designs the state space to handle limited observability by incorporating differences rather than absolute measurements of hard-to-measure variables like OD demand.

Real-world limited observability: sk = [ni,k, ∆ni,k, ∆2 ni,k, Mij,k] (Eq. 9)

]

The improved AI system can achieve:

  • Practical Real-World Deployment: The algorithm maintains performance even when critical data points (like real-time OD demand estimates) are unavailable or noisy. It learns to make optimal decisions based on the temporal evolution of known, observable quantities (differences and accumulation rates).

  • Trade-off Optimization: The system can operate effectively within realistic constraints, achieving a quantifiable performance gain (e.g., 4.8% average improvement in TTS under limited observability) by intelligently trading off perfect information for operational feasibility.

)4. Superior Performance Under Adversarial Conditions

Empirical results show that the proposed antifragile RL algorithm significantly outperforms traditional benchmarks (MPC, MPC-MHE, baseline RL) when facing increasing demand and supply disruptions, both in full and limited observability settings.

the proposed algorithm exhibits both superior performance and reduced variance (Section 6.1)

ultimate performance improvement of 41.9% under maximal demand/supply disruptions (Section 7)

]

The improved AI system can achieve:

  • Disruption Resilience: It demonstrates superior resilience, achieving gains up to 41.9% over baselines when facing combined demand and supply shocks. This means the system not only recovers from a disruption but actively learns to utilize the disorder (the disruption) to its advantage, leading to sustained performance gains under stress.

In summary, this research transforms standard Reinforcement Learning for traffic control into an Antifragile RL framework that is:

  1. More predictive (via derivatives).

  2. More stable (via damping).

  3. More adaptive (via redundancy).

  4. More resilient (quantified via skewness analysis), resulting in a system that gains performance and robustness when facing the exact types of unpredictable disruptions it was designed to anticipate and thrive within.

Abstract

The optimal operation of transportation systems is often susceptible to unexpected disruptions. Many established control strategies reliant on mathematical models can struggle with real-world disruptions, leading to significant divergence from their anticipated efficiency. This study integrates the cutting-edge concept of antifragility with learning-based traffic control strategies to optimize urban road network operations under disruptions. Antifragile systems not only withstand and recover from stressors but also thrive and enhance performance in the presence of such adversarial events. Incorporating antifragile modules composed of traffic state derivatives and redundancy, a deep reinforcement learning algorithm is developed. Subsequently, it is evaluated in a cordon-shaped transportation network and a case study with real-world data. Promising results highlight that the proposed algorithm provides: (i) superior performance achieving up to 27.6% and 41.9% performance gain over baselines under increasing demand and supply disruptions, (ii) lower distribution skewness under disruptions, demonstrating its relative antifragility against baselines, (iii) effectiveness under limited observability due to real-world data availability constraints, and (iv) the robustness and transferability to be combined with various state-of-the-art RL frameworks. The proposed antifragile methodology is generalizable and holds potential for applications beyond traffic engineering, offering integration into control systems exposed to disruptions across various disciplines.

Sources

Related papers