Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing

arXiv:2603.28900 · cs.RO, cs.AI, cs.LG, cs.SY, eess.SY · Submitted 2026-03-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing".

Jane: The paper was written by Alex Zongo, Filippos Fotiadis, Ufuk Topcu and Peng Wei from Department of Mechanical and Aerospace Engineering, George Washington University and Oden Institute for Computational Engineering & Sciences, University of Texas at Austin.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary and Core Mechanics: Tom: We’ve seen the problem, but how does the robust policy actually handle this corrupted information? It's a complex multi-agent system.

Jane: They use Multi-Agent Reinforcement Learning, or MARL, where each drone uses its local observation to make decisions about speed and path planning.

Lu: The complexity is that they don’t just assume the AI is smart; they build a specific mathematical framework around the uncertainty itself.

Meng: The core of this system seems to be that when the GPS gives you bad data, your entire perceived traffic state—your own position plus all nearby drones—might be wrong.

Lalam: That's the "correlated full-state corruption" that makes this so difficult; it’s not just one bad reading, it’s a whole distorted picture of the airspace.

Tom: And to handle this, they integrate their closed-form adversarial solution into a standard PPO algorithm.

Jane: It uses PPO—Proximal Policy Optimization—to ensure that even when the input data is contaminated, the learning process doesn't jump wildly or forget what it learned before this corruption exists.

Lu: The key insight here is using this analytical worst-case prediction to guide the training, rather than just letting the AI stumble upon robustness through many randomized trials.

Meng: From an engineering view, this means we can deploy a robust policy that knows precisely how much "noise" it needs to ignore in order to maintain safe separation.

Lalam: The system is designed not just to survive corruption but to maintain high performance across a predictable spectrum of degradation, which is crucial for trust.

The Robustness Approach: Tom: Let's talk about the improvements, specifically the way they make the AI robust through regularization. It sounds like there are two different ways they stabilize it.

Jane: They introduce two KL-based regularizers that act almost like a safety net for the policy itself.

Lu: One of them ensures that even if you are looking at corrupted data, your decision-making process remains fundamentally similar to the original, clean version of the policy.

Meng: And the second one acts as an anchor—it keeps the trained AI close to a "nominal" or pre-trained stable version of itself.

Lalam: That combination is brilliant because it allows us to achieve robustness without sacrificing performance; we can be safe without being inefficient.

Tom: So, if that invariance regularization works, it directly bounds the expected degradation in decision-making.

Jane: That’s right, Proposition one shows that this regularization directly limits how much worse the policy will perform under corruption.

Lu: It prevents the AI from "hallucinating" a safe path when its perception is compromised, which is a huge step forward for safety standards.

Meng: I appreciate the clear separation of concerns; you are ensuring both consistency and stability simultaneously, which makes this highly dependable in practice.

Lalam: This allows us to move beyond just surviving failures toward achieving predictable, stable performance in a complex urban environment.

Conclusion and Impact: Tom: We’ve covered the technical groundwork, but let's bring it back to the big picture—the real-world results of "Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing."

Jane: The simulation results are incredibly encouraging, showing that even when GPS is corrupted up to thirty-five percent, the robust policy maintains near-zero collisions.

Lu: It really demonstrates that we can handle adversarial conditions without needing a massive computational overhead or an endless cycle of adversarial training.

Meng: From an operational standpoint, seeing this level of performance at thirty-five percent corruption suggests that real-world deployment is much closer than I previously thought.

Lalam: The ultimate impact, Lalam feels, is that the trust we put in autonomous systems can be drastically increased because we have a formal mathematical guarantee of how well they will perform under attack.

Tom: It’s clear that this research has provided a robust framework for autonomous navigation in challenging urban airspace.

Jane: Indeed, and as we conclude this discussion on "Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing," it shows the power of combining analytical math with advanced AI techniques.

Lu: It's a beautiful synthesis of theoretical rigor and practical application, showing what's possible when you attack the problem from a different angle.

Meng: I think this paper sets a new standard for reliability in safety-critical AI systems.

Lalam: The advancement provides not just better drones, but better standards for the trust we place in technology itself.

Conclusion: Tom: It's clear that we’ve successfully navigated the complexities of this research, moving from a theoretical problem to real-world performance.

Jane: And that's what makes the results so exciting; we have seen how these autonomous systems can maintain safety even when they are being deliberately misled.

Lu: The shift in modeling the environment as a zero-sum game is a huge paradigm change for me, because it means we are no longer just hoping for random robustness.

Meng: That analytical approach allows us to build systems that will actually work in unpredictable urban environments, which is what we really need when deployment becomes practical.

Lalam: The ability to quantify risk so precisely as the authors have done fundamentally changes how society perceives and trusts autonomous technology.

Tom: It certainly elevates the conversation from simple "fail-safes" to true "resilience," which is a major accomplishment in itself.

Jane: And it's a perfect example of combining high-level mathematical rigor with practical, scalable AI techniques.

Lu: I think this methodology opens up so many new avenues for subsequent researchers who will be tackling even more complex traffic scenarios.

Meng: It gives us a clear roadmap for engineering that suggests we can build systems that handle the worst-case scenario without the massive overhead of endless adversarial training.

Lalam: The advancement provides not just better drones, but a new standard of reliability for all the autonomous services we are building.

Tom: We really think this research, "Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing," sets a powerful new benchmark.

Jane: It’s been fascinating to see how far the field has come in addressing these critical issues of trust and safety.

Lu: I just hope we get to discuss more of the theoretical implications next time around, because there is so much more to explore.

Meng: We certainly will; for now, I'm just happy that this was a problem we can solve with real-world reliability in mind.

Alex Zongo, Filippos Fotiadis, Ufuk Topcu, Peng Wei

Department of Mechanical and Aerospace Engineering, George Washington University · Oden Institute for Computational Engineering & Sciences, University of Texas at Austin

cs.RO, cs.AI, cs.LG, cs.SY, eess.SY

Submitted: 2026-03-30

Updated: 2026-08-29

Importance score: 87/100

The gist: The paper, "Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing," presents a framework for ensuring safe separation of small Unmanned

Key concepts

Multi-Agent Reinforcement Learning (MARL)
MARL is the core system where multiple autonomous agents, or drones, operate. Each drone uses its local observations to make decisions regarding its specific speed and path planning within the airspace.
Robust Policy
This refers to the AI's ability to function reliably despite bad data. It is achieved using KL-based regularizers that act as a safety net, ensuring that even if input data is corrupted, the decision-making process remains fundamentally consistent with its original performance.
PPO (Proximal Policy Optimization)
PPO is the standard algorithm integrated into this system. It ensures that when the input data is contaminated or corrupted, the learning process does not jump wildly or forget previously learned safe navigation methods.

Terminology

Summary

The paper, Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing, presents a framework for ensuring safe separation of small Unmanned Aircraft Systems (sUAS) when their Global Positioning System (GPS) data is corrupted by degradation or adversarial spoofing.

Problem Formulation and Motivation

The paper addresses the challenge that in cooperative surveillance, sUAS rely on GPS-derived position broadcasts. When these broadcasts are corrupted, the entire observed air traffic state becomes unreliable. The authors model this situation as a zero-sum game between the agents and an adversary who perturbs the observed state to maximize safety degradation.

The problem is formulated as a Markov Decision Process (MDP) with adversarially corrupted observations. The true local traffic state S t is represented by, for up to m nearby intruders, s in R 1 times m+1. However, the observation available to the agent at time t, denoted t, may be corrupted.

The authors model GPS corruption using R-contamination uncertainty. Conditional on the true next state S t+1, the observation t+1 is generated as:

OR p S t+1 q = (1 - R) delta S t+1 + R delta t (3)

where R is the corruption probability. The adversarial observation t belongs to a state-dependent uncertainty set:

S t+1 q =: - S t+1 kappa u (5)

The reward r p S t+1, u q encourages safe separation while penalizing control effort:

r p S t+1, u q = -sum j=1 m [q doj p S t+1 q - g puq] (6)

Core Contributions and Theoretical Guarantees

The paper makes several key theoretical contributions:

  1. Closed-Form Adversarial Perturbation: The authors derived a closed-form approximate solution for the worst-case adversarial perturbation, bypassing the need for extensive adversarial training.

t, FO = kappa d sign grad V pi p S t+1 (16)

  1. Second-Order Accuracy: They proved that this first-order (FO) approximation approximates the true worst-case perturbation with an error that is second-order in the perturbation magnitude:

V pi p q - V pi p S t+1 about kappa d grad V pi p S t+1 + O(kappa 2) (20)

  1. Performance Bounding: They proved that Kullback-Leibler (KL)-based policy regularization bounds the expected performance loss induced by corrupted observations, yielding a principled robustness performance tradeoff.

E pi V pi p S q - V corrupt p S q Q squared B (32)

Robust Policy Optimization Method

The authors solve the robust objective using a proximal policy optimization (PPO) approach with a shared actor-critic network. The training involves two phases: pretraining under clean observations (R=0), followed by robust training under R-contaminated observations.

To achieve robustness, the method incorporates two KL regularizers:

  1. Perturbation-Invariance Regularizer (Linv): This term penalizes the policy's sensitivity to observation corruption, encouraging similar decisions under clean and perturbed inputs:

Linv p theta = E S q D KL pi theta(S q) pi theta(q) (29)

  1. Teacher-Anchor Regularizer (Lanchor): This term keeps the learned policy close to the pretrained nominal teacher policy theta on clean observations:

Lanchor p theta = E S q D KL pi(S q) pi theta(S q) (30)

The complete robust actor objective is defined as:

L actor p theta = L clip p theta + lambda inv Linv p theta + lambda anchor L anchor, (31)

Experimental Evaluation and Results

The framework was evaluated in a high-density sUAS simulation using the BlueSky air traffic simulator, operating within a structured en-route airspace. The comparison was made between the nominal policy (trained without adversarial corruption and R=0) and the full robust policy.

The results demonstrate significant superiority of the robust approach:

  • Performance under Corruption: In high-density sUAS simulation, we observe near-zero collision rates under corruption levels up to 35%, outperforming a baseline policy trained without adversarial perturbations. (Abstract)

  • Safety Metrics: Figure 3 shows that as R increases, the nominal policy's near-zero NMAC count deteriorates sharply after R about 0.25. In contrast, the robust policy sustains near-zero NMACs through R 0.35 and degrades more gracefully beyond.

  • Ablation Study: The results of Figure 5 show that invariance regularization without anchoring destabilizes training, while anchoring alone degrades sharply at high R. The full method combines both to achieve the strongest performance.

The conclusion states that the simulation experiments demonstrate the effectiveness of the robust MARL policy and its superiority compared to a baseline policy, on maintaining safe separation between the sUASs under GPS degradation and spoofing.

Improvements for AI systems

Based on a rigorous analysis of the provided scientific paper, here are the specific technical improvements that can be applied to improve AI systems, along with a detailed description of what these improved systems can achieve.


The core innovations presented in this work move beyond merely testing robustness; they provide mathematically rigorous, tractable methods for engineering it. These improvements generalize far beyond small Unmanned Aircraft Systems (sUAS) to any decentralized system reliant on potentially corrupted sensor data (e.g., autonomous vehicles, industrial robotics).

The Improvement: The paper derives a closed-form expression for the worst-case adversarial perturbation (FO) by utilizing a first-order Taylor expansion of the value function (V). This replaces computationally expensive, iterative adversarial training with a direct analytical calculation.

  • Key Mechanism: FO = St' - kappa times sgn(grad V pi(St'))

  • Impact on AI System: Computational Efficiency. The system can instantly calculate the most detrimental possible input (the adversarial state) without running repeated inner minimization loops, enabling real-time, high-frequency policy updates and decision-making.

The Improvement: A formal proof is provided demonstrating that the closed-form approximation of the worst-case perturbation is accurate to a second order (O(kappa 2) relative to the corruption magnitude kappa).

  • Key Mechanism: The error gap between the exact worst-case value and its bounded estimate is quantified by O(kappa 2) / L V times kappa squared.

  • Impact on AI System: Trust and Calibration. This allows system designers to precisely quantify the theoretical limits of their model. It provides a rigorous guarantee that the approximation used in practice (the closed-form solution) is highly reliable, mitigating uncertainty regarding how much performance is being lost to approximation error.

The Improvement: Utilizing Kullback-Leibler (D KL) regularization to explicitly bound the expected performance loss between a clean observation and an adversarial observation.

  • Key Mechanism: Proposition 1 proves that controlling the expected D KL difference limits the one-step value degradation: E[V] 2Q times

sqrt ES[D KL.

  • Impact on AI System: Proactive Safety Assurance. The system doesn's merely react to a collision; it is trained to maintain a quantifiable, predictable level of performance degradation under attack. This allows safety engineers to set explicit thresholds (e.g., Performance cannot drop more than X% regardless of input corruption) rather than relying on empirical testing.

The Improvement: Integrating two distinct KL-based regularizers into the PPO objective:

  • Invariance (Linv): Enforcing similarity between the policy's output on clean and adversarial inputs, D KL(pi(S) pi(FO)).

  • Anchoring (Lanchor): Keeping the learned robust policy close to a pre-trained nominal teacher policy, D KL(pi teacher(S) pi robust(S)).

  • Impact on AI System: Robust Training Stability. This prevents the catastrophic forgetting or over-optimization common in robust RL. The system is forced to learn robustness without sacrificing its original, effective performance in benign environments.


By implementing these technical improvements, a decentralized, multi-agent AI system (applicable to autonomous vehicles or drone swarms) will possess the following capabilities:

1. Guaranteed Safety Under Attack:

The system can operate in environments where sensor data is deliberately manipulated (e.g., GPS spoofing, false radar returns). It maintains a provable level of safety, ensuring that the probability of a critical failure (e.g., collision, loss of separation) remains below a predefined tolerance, even when facing an adversary who knows its training methodology.

2. Real-Time Robust Decision Making:

The system can execute complex, multi-agent maneuvers in real-time because it does not need to perform computationally expensive iterative adversarial search for the worst case. It simply uses the analytical closed-form solution (FO) to determine the safest action based on its current perception of a corrupted environment.

3. Adaptive Performance Degradation:

The system can be deployed in safety-critical scenarios where perfect performance is impossible (e.g., extreme weather, heavy spoofing). Instead of failing outright, the system gracefully degrades its operational envelope—slow down or increase separation margins—in direct proportion to the level of detected sensor corruption, as predicted by its KL-based bounds.

4. Certifiable Robustness:

The resulting AI system provides a safety certificate. A manufacturer can provide proof that the policy's performance degradation is mathematically bounded by O(kappa 2) and D KL limits, offering a level of accountability and reliability unattainable with traditional empirical testing alone.

Sources

Related papers