Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization

arXiv:2511.20718 · cs.LG, cs.AI, cs.CL · Submitted 2025-11-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization".

Tom: Reinforcement learning (RL) algorithms like PPO and GRPO often suffer from unstable optimization dynamics when applied to training large language models (LLMs) for multi-turn agentic tasks in off-policy settings.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: We just touched on the setup, so what is the central idea behind this SORL framework? What exactly are they proposing to fix these issues with?

Jane: The core thesis of this paper is to introduce SORL, which uses two specific mechanisms: turn-level importance sampling and clipping-triggered normalization. The goal is to align the policy optimization process with the inherent structure of multi-turn reasoning while also adaptively suppressing those noisy off-policy updates.

Lu: I see how turn-level importance sampling tries to solve that granularity mismatch by defining weights based on what happens before a certain turn, which should aggregate token advantages into a meaningful signal for the entire turn rather than treating every single token as an isolated event.

Meng: So, if they are aggregating advantages at the turn level, does that mean we're moving away from optimizing every single word individually? That seems like it could simplify how we think about credit assignment in long sequences.

Lalam: If the system is more structured around turns, I think our models will be better at maintaining context and coherence across longer interactions, which is something I really value for complex tasks.

Tom: Exactly, and then they pair that with clipping-triggered normalization to handle the variance spikes caused by aggressive updates or training-inference mismatches in off-policy settings. That combination seems designed to keep the learning process smooth rather than volatile.

Jane: Right, so they introduce a turn-level PPO objective where the importance sampling weights are defined specifically to achieve this turn-wise aggregation of token advantages, which should give us a more stable policy gradient computed at the level of turns instead of individual tokens.

Lu: It’s like shifting the focus from optimizing every single pixel in an image to optimizing larger regions, which is a common technique in computer vision and makes intuitive sense for sequential reasoning.

Meng: I’m still thinking about how they define that clipping-triggered normalization; how does it specifically correct the bias introduced by those clipping signals when we're dealing with highly variable importance sampling ratios? That part sounds technically tricky to implement correctly.

Lalam: If it helps suppress those unreliable updates, then any improvement in stability translates directly into higher quality outputs for the agentic tasks we are working on.

Conclusion: Tom: So we've walked through how SORL tackles the core problems of granularity mismatch and update variance in off-policy training for LLM agents. What does this paper, published by Li et al., actually mean for the future of agentic AI?

Jane: The main implication is that we can develop more robust training pipelines for long-horizon tasks without needing constant, tedious manual tuning of hyperparameters just to keep the optimization stable. They show that this structured approach leads to more conservative and reliable learning dynamics across various benchmarks.

Lu: From a theoretical standpoint, it provides a principled way to connect the local token optimization with the global structure of reasoning steps in an agentic workflow, which opens up new avenues for how we design these sequential models.

Meng: Practically speaking, if this framework consistently prevents performance collapses that we see in standard PPO or GRPO when dealing with complex agentic workflows, it means we can deploy these agents with a much higher degree of confidence in their reliability.

Lalam: For me, this means the AI systems we build will become much more trustworthy because the underlying training process is fundamentally sounder and less susceptible to those unpredictable failures that undermine user trust.

Tom: So, to wrap up, SORL gives us a unified framework that combines turn-level credit assignment with clipping-triggered normalization. It moves the needle from unstable optimization dynamics to more predictable learning outcomes for these complex LLM agents in off-policy settings.

Jane: Precisely; it establishes a principled way forward by addressing those two root causes directly, yielding better performance without requiring us to rely on those heuristic tuning tricks we usually have to resort to.

Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao

Texas A&M University

cs.LG, cs.AI, cs.CL

Submitted: 2025-11-25

Updated: 2026-10-05

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: Reinforcement learning (RL) algorithms like PPO and GRPO often suffer from unstable optimization dynamics when applied to training large language models (LLMs) for multi-turn agentic tasks in

Key concepts

Turn-Level Importance Sampling
This technique adjusts the policy update by weighting tokens based on their relevance within a specific turn of conversation. It aligns the token-level optimization with the natural flow of multi-turn reasoning, ensuring that credit is assigned coherently across turns rather than just individual words.
Clipping-Triggered Normalization
This mechanism stabilizes updates by explicitly correcting bias introduced by PPO's clipping when using off-policy data. It decomposes the gradient to separate standard policy changes from clipping effects, allowing the framework to adaptively reduce step size when off-policy variance becomes too high.
Granularity Mismatch
This is an instability where the optimization happens at a fine level (individual tokens), but the task structure is coarse (multi-turn interactions). SORL addresses this by aggregating token advantages into turn-level signals, ensuring that policy gradients reflect meaningful turn-based reasoning rather than noisy token details.

Terminology

Summary

Reinforcement learning (RL) algorithms like PPO and GRPO often suffer from unstable optimization dynamics when applied to training large language models (LLMs) for multi-turn agentic tasks in off-policy settings. This work proposes SORL, a framework that stabilizes these training processes by introducing turn-level importance sampling and clipping-triggered normalization, which addresses the granularity mismatch between token-level optimization and turn structure and suppresses high-variance off-policy updates.

How it works

The core of the proposed SORL framework involves two key components designed to align policy optimization with multi-turn structures and adaptively suppress unreliable updates. First, turn-level importance sampling is introduced to handle the granularity mismatch between token-level policy optimization and turn-structured interactions. This mechanism aligns policy optimization with the natural turn structure of multi-turn reasoning, enabling structure-aware credit assignment that balances the noise of token-level updates and the coarseness of sequence-level objectives. This is formalized by defining a turn-level PPO objective where importance sampling weights are defined as:

w turn k(θ):= πθ(y k x, y<k) / πθold(y k x, y<k).

The lemma proves that this approach induces a turn-wise aggregation of token-level advantages, meaning all tokens within the same turn share a common, length-normalized credit signal, which results in policy gradients computed at the granularity of turns rather than individual tokens.

How it works (Continued)

Second, clipping-triggered normalization stabilizes increasingly off-policy updates by explicitly correcting clipping-induced bias in PPO gradients. This mechanism targets the high variance induced by off-policy importance sampling ratios that become highly variable due to aggressive updates and factors such as training–inference mismatch. The framework decomposes the PPO gradient into a standard policy gradient term and a clipping bias term, C(θ), which captures the contribution of tokens for which clipping is active. To mitigate this, a surrogate gradient estimator is proposed:

∇θJSO-PPO(θ):= 1 / Cturn(θ)2 ∇θJTurn-PPO(θ).

Where Cturn(θ) is defined as Eτ∼πθ [Cturn(θ; τ)], and the stabilized policy gradient is rescaled by the inverse of this norm, adaptively reducing the effective step size when off-policy variance dominates.

Key Contributions

The paper identifies two fundamental sources of instability in off-policy training for LLM agents: (i) a granularity mismatch between token-level optimization and turn-level interactions, and (ii) the accumulation of variance from unreliable off-policy samples, where state–action pairs are poorly evaluated. The primary contributions are:

  1. Empirical diagnosis of these two root causes unique to the agentic setting.

  2. Proposal of a unified off-policy reinforcement learning framework that integrates turn-level credit assignment with a clipping-triggered normalization mechanism to suppress unreliable updates, yielding more conservative and robust learning dynamics.

  3. Demonstration that the framework is algorithm-agnostic and can be instantiated in other methods beyond PPO, such as GRPO, where SO-GRPO achieves substantially improved training stability over standard GRPO.

Experimental Validation

The proposed SORL framework was evaluated on multi-turn search benchmarks, including general question answering and medical multiple-choice QA tasks. Experimental results show that both SO-PPO and SO-GRPO consistently prevent training instabilities and performance collapses observed in standard PPO and GRPO, maintain lower clipping ratios, and achieve superior or comparable task performance. Ablation studies confirm that the method remains stable under increasing off-policyness, demonstrating that the turn-level importance sampling aligns credit assignment with reasoning-search interactions, while clipping-triggered normalization adaptively scales gradient updates to suppress destabilizing gradients. Furthermore, SO-GRPO was shown to further outperform GSPO on the NQ dataset.

Conclusion

SORL establishes a principled framework for stabilizing off-policy reinforcement learning in multi-turn LLM agent training by combining turn-level credit assignment and clipping-triggered normalization. This approach successfully mitigates the granularity mismatch and adapts to increasing off-policyness, resulting in more stable optimization dynamics without requiring heuristic tuning or early stopping, leading to improved task performance across various benchmarks.

The gist: SORL stabilizes off-policy reinforcement learning for long-horizon LLM agent training by combining turn-level importance sampling and clipping-triggered normalization to address granularity mismatch and high-variance updates.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems by implementing the SORL framework, along with what those improved systems will be capable of:


The proposed SORL (Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Training) framework introduces two core stabilizing mechanisms: turn-level credit assignment and clipping-triggered normalization. Applying this to existing LLM RL pipelines (like PPO and GRPO) yields the following specific improvements:

  1. The ability to train LLM agents for multi-turn, long-horizon tasks without catastrophic performance collapse or requiring manual early stopping heuristics.

  2. The development of more robust and reliable agentic decision-making by ensuring that policy updates are not dominated by high-variance, unreliable off-policy samples.

Specific Capabilities of the Improved AI Systems:

  1. A new class of LLM agents trained via SO-PPO (and SO-GRPO) will be capable of performing complex, multi-step reasoning tasks (e.g., multi-hop QA, complex problem solving) with significantly higher and more consistent success rates compared to standard PPO/GRPO baselines.

  2. The system will exhibit superior stability during long training runs under increasingly off-policy data regimes (i.e., as mini-batch reuse increases), maintaining lower KL divergence and more reliable performance trajectories without the need for careful checkpoint selection or heuristic tuning of clipping ratios.

  3. The agents will demonstrate better coordination between internal reasoning steps and external tool use (retrieval/search). The turn-level importance sampling ensures that credit assignment is correctly attributed to specific reasoning phases (e.g., analysis, query formulation, information processing) rather than being diluted by token-level noise, leading to more effective and less erratic tool invocation strategies.

  4. The systems will show stronger generalization across diverse domains (like medical QA) because the normalization mechanism adaptively downweights updates based on the reliability of the sampled data batch, preventing reward hacking or over-optimization on non-semantic tokens.

  5. The resulting agents will possess a more conservative and robust policy, making them suitable for high-stakes applications where training stability is paramount (e.g., medical diagnostics, technical troubleshooting), as evidenced by their ability to maintain steady performance improvements throughout the entire optimization horizon.

Abstract

Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing. SORL introduces mechanisms that align policy optimization with the structure of multi- turn interactions and adaptively suppress unreliable off-policy updates, yielding more conserva- tive and robust learning dynamics. Within this framework, we instantiate two stabilized algo- rithms: SO-PPO and SO-GRPO. Both algorithms are designed to mitigate gradient variance and prevent optimization collapse without requiring careful early stopping or heuristic tuning. We evaluate SO-PPO and SO-GRPO on benchmarks spanning open-domain QA, multi-hop QA, and medical multiple-choice QA, and further assess their transfer to asynchronous RL for mathe- matical reasoning by training on DAPO-Math-17k and validating on AIME-2024. These results demonstrate that SORL provides a practical, scalable, and general framework for stabilizing re- inforcement learning in multi-turn LLM agent training and asynchronous RL for mathematical reasoning.

Sources

Related papers