Learning Perturbation Robust Policies for LLM Agents with Stable Optimization

arXiv:2609.34064 · cs.LG, cs.AI · Submitted 2026-09-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Learning Perturbation Robust Policies for LLM Agents with Stable Optimization".

Jane: Reinforcement learning policies trained for long-horizon large language model agents are sensitive to various policy perturbations, such as hidden state noise, pruning, and quantization.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, wrapping up our discussion on "Learning Perturbation Robust Policies for LLM Agents with Stable Optimization," the authors have really laid out a path for making these long-horizon agents more resilient to noise and pruning during training.

Jane: They've shown that by studying how policy updates behave when perturbed, they can propose SPrPO, which uses adaptive perturbations guided by estimated improvement and sensitivity to control the process <ref:2609.34064#pg1>.

Lu: The theoretical extension of monotonic improvement analysis to this perturbation setting is significant because it gives us a formal condition for when an update will be stable <ref:2609.34064#pg2>.

Meng: From a practical viewpoint, the ability to systematically handle different types of perturbations like quantization and pruning with this controlled method means we can deploy LLM agents with a higher degree of confidence in their reliability <ref:2609.34064#pg1>.

Lalam: For me, the implication is that we can build AI systems that are fundamentally more dependable, not just performant on clean data, but resilient even when the underlying model or its environment experiences some form of noise or corruption <ref:2609.34064#pg1>.

Tom: It's a method for controlling the perturbation scheme during RL training to improve robustness while keeping optimization stable, which they validated across many benchmarks like ALFWorld and WebShop <ref:2609.34064#pg0>.

Jane: The title itself speaks to the core achievement: learning policies that are robust specifically through this perturbation-aware optimization process, rather than just training on perfect data.

Lu: This work pushes the understanding of how to train complex, sequential decision-making agents in a way that respects their sensitivity to internal changes <ref:2609.34064#pg1>.

Meng: I think this points toward a future where we don't just optimize for peak performance on training sets, but we optimize for reliable operation in unpredictable, noisy real-world environments <ref:2609.34064#pg1>.

Lalam: It means the next generation of AI agents will be less fragile and more capable of handling the messy reality of deployment <ref:2609.34064#pg1>.

Conclusion: Segment: Conclusion — Title and Implications**

Tom: So, we've covered how SPrPO tackles those nasty perturbations like noise and pruning during LLM agent training, and now we need to talk about what that title actually means for us as listeners.

Jane: Yeah, the paper is called "Learning Perturbation Robust Policies for LLM Agents with Stable Optimization," which basically tells us they're finding a way to train these complex AI agents so they don't break when things get a little messy during the learning process.

Lu: It’s fascinating because it moves past just getting high scores on clean data; it addresses the instability that creeps in when you actually try to deploy these long-horizon models in the real world, where everything is inherently noisy.

Meng: From my side, I'm thinking about how this stability translates into reliability; if we can keep the training process stable even with these perturbations, it suggests a more predictable path to building production-ready AI agents that don't fail unexpectedly.

Lalam: For me, the implication is huge because it suggests we can build AI systems that aren't just smart on paper but are resilient in practice, which could fundamentally improve how we trust and use these powerful tools in our daily lives.

Tom: Exactly! It’s not about making a slightly smarter model; it’s about making a more dependable one through this specific optimization technique.

Jane: The authors essentially show us that by being careful about how we perturb the system during training, we can achieve policies that are much harder to break when they encounter real-world noise later on.

Lu: I think the core idea is establishing a theoretical bridge between those classical optimization methods and this more complicated reality of adversarial perturbations.

Meng: It’s a practical win because it gives us an actionable strategy for tuning our RL training pipelines so that we aren't constantly fighting against instability as we scale up.

Lalam: I really hope this research means future AI agents can be deployed in mission-critical systems where even minor noise could cause major failures, and this work offers a way to mitigate that risk significantly.

Tom: This whole concept of "perturbation robustness" is really something to chew on for anyone building these kinds of complex AI agents, and it opens up a lot of avenues for future research.

Department of Electrical and Computer Engineering University of Arizona

cs.LG, cs.AI

Submitted: 2026-09-28

Updated: 2026-09-28

Importance score: 92/100

The gist: Reinforcement learning policies trained for long-horizon large language model agents are sensitive to various policy perturbations, such as hidden state noise, pruning, and quantization.

Key concepts

Perturbation Robust Policy
A policy is considered perturbation robust if its performance remains stable even when the underlying model or environment is slightly altered by noise, pruning, or quantization. The paper analyzes the mathematical conditions required for a policy update to maintain this stability despite these external disturbances.
Monotonic Improvement Analysis (TRPO Extension)
This is a theoretical framework used to ensure that every policy update leads to an improvement in performance. The authors extend the traditional Trust Region Policy Optimization (TRPO) analysis into the perturbation setting, establishing a sufficient condition for stable improvement when perturbations are present.
SPrPO Algorithm
SPrPO is the practical method introduced to achieve robustness. It dynamically adjusts the scale of Gaussian noise added to the model's hidden states based on estimated expected improvement and channel sensitivity. This adaptive noise scaling ensures that training remains stable while actively mitigating performance degradation caused by perturbations.

Terminology

Summary

Reinforcement learning policies trained for long-horizon large language model agents are sensitive to various policy perturbations, such as hidden state noise, pruning, and quantization. This work introduces Stable Perturbation-Robust Policy Optimization (SPrPO), a perturbation-based method designed to improve the robustness of these policies during training while preserving stable optimization.

How it works

The core idea involves studying how policy updates preserve stable monotonic improvement when subjected to perturbations. The authors first introduce the notion of a perturbation robust policy and analyze conditions under which perturbed policy updates maintain this stability. Based on this analysis, they propose SPrPO, which applies adaptive and sensitivity-aware perturbations during RL training. This method is guided by a practical condition for controlling perturbations during optimization to ensure stability.

Theoretical Foundation

The theoretical analysis extends the classical monotonic improvement analysis of TRPO to the perturbation setting. They establish a sufficient condition for stable policy improvement and characterize the limit points of resulting updates as being perturbation-robust. Specifically, they derive Corollary 1, which provides a tractable characterization for Gaussian perturbations under suitable regularity conditions. This corollary states that if a specific stability condition involving expected improvement, policy drift, and perturbation-induced variance is met (Equation 8), then the probability of positive perturbed improvement is greater than or equal to 1 minus β.

Practical Implementation (SPrPO)

SPrPO is realized through a practical algorithm summarized in Algorithm 1. The method applies controlled Gaussian perturbations in the hidden space of the FFN blocks. The noise scale for each channel, denoted as σk,lc, is dynamically adjusted based on two factors: the overall perturbation magnitude ρ and a sensitivity exponent α. Specifically, the standard deviation of channel perturbation to increase with the estimated improvement and decrease with channel sensitivity. The algorithm estimates these quantities by reusing the rollout batch from the current policy update to estimate channel sensitivity (Sbk+1,lc) and expected improvement (mbk+1).

Evaluation and Results

The effectiveness of SPrPO is validated on long-horizon benchmarks like ALFWorld and WebShop across multiple perturbation types, including Gaussian noise, dropout, structured pruning, unstructured pruning, and quantization. Empirical results show that SPrPO exhibits more stable optimization during training and better preserves task performance under perturbation. For instance, in the evaluation of ALFWorld on WebShop under structured pruning (r=0.20), SPrPO achieved a success rate of 89.17%, compared to 41.73% for Vanilla, demonstrating significantly improved robustness against this perturbation type. Qualitative examples further illustrate this by showing SPrPO's ability to successfully navigate malformed action tags and complete multi-step tasks where Vanilla fails due to interface errors or unproductive action repetition.

Key Contributions

The main contributions of this work are:

  1. Being the first to study perturbation robustness for RL-trained LLM agents and formulate the problem under a general distribution of policy perturbations.

  2. Theoretically extending the monotonic policy improvement analysis of TRPO to the perturbation setting and characterizing a sufficient condition for stable policy improvement, establishing that limit points are perturbation-robust.

  3. Deriving a practical condition for controlling perturbations during optimization and developing SPrPO.

  4. Demonstrating that SPrPO improves robustness across diverse policy perturbations while maintaining training stability, as evidenced by the experimental results in Tables 5, 6, and 7.

The gist: Stable Perturbation-Robust Policy Optimization (SPrPO) is a perturbation-based method for RL training that improves the robustness of LLM agents while preserving training stability.

Improvements for AI systems

Based on the scientific paper LEARNING PERTURBATION ROBUST POLICIES FOR LLM AGENTS WITH STABLE OPTIMIZATION, here are specific, high-impact improvements that can be implemented in AI systems, along with the resulting capabilities:


)Improvements and Enabled Capabilities:

Sources

Related papers