Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization".
Tom: Reinforcement learning (RL) algorithms like PPO and GRPO often suffer from unstable optimization dynamics when applied to training large language models (LLMs) for multi-turn agentic tasks in off-policy settings.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: We just touched on the setup, so what is the central idea behind this SORL framework? What exactly are they proposing to fix these issues with?
Jane: The core thesis of this paper is to introduce SORL, which uses two specific mechanisms: turn-level importance sampling and clipping-triggered normalization. The goal is to align the policy optimization process with the inherent structure of multi-turn reasoning while also adaptively suppressing those noisy off-policy updates.
Lu: I see how turn-level importance sampling tries to solve that granularity mismatch by defining weights based on what happens before a certain turn, which should aggregate token advantages into a meaningful signal for the entire turn rather than treating every single token as an isolated event.
Meng: So, if they are aggregating advantages at the turn level, does that mean we're moving away from optimizing every single word individually? That seems like it could simplify how we think about credit assignment in long sequences.
Lalam: If the system is more structured around turns, I think our models will be better at maintaining context and coherence across longer interactions, which is something I really value for complex tasks.
Tom: Exactly, and then they pair that with clipping-triggered normalization to handle the variance spikes caused by aggressive updates or training-inference mismatches in off-policy settings. That combination seems designed to keep the learning process smooth rather than volatile.
Jane: Right, so they introduce a turn-level PPO objective where the importance sampling weights are defined specifically to achieve this turn-wise aggregation of token advantages, which should give us a more stable policy gradient computed at the level of turns instead of individual tokens.
Lu: It’s like shifting the focus from optimizing every single pixel in an image to optimizing larger regions, which is a common technique in computer vision and makes intuitive sense for sequential reasoning.
Meng: I’m still thinking about how they define that clipping-triggered normalization; how does it specifically correct the bias introduced by those clipping signals when we're dealing with highly variable importance sampling ratios? That part sounds technically tricky to implement correctly.
Lalam: If it helps suppress those unreliable updates, then any improvement in stability translates directly into higher quality outputs for the agentic tasks we are working on.
Conclusion: Tom: So we've walked through how SORL tackles the core problems of granularity mismatch and update variance in off-policy training for LLM agents. What does this paper, published by Li et al., actually mean for the future of agentic AI?
Jane: The main implication is that we can develop more robust training pipelines for long-horizon tasks without needing constant, tedious manual tuning of hyperparameters just to keep the optimization stable. They show that this structured approach leads to more conservative and reliable learning dynamics across various benchmarks.
Lu: From a theoretical standpoint, it provides a principled way to connect the local token optimization with the global structure of reasoning steps in an agentic workflow, which opens up new avenues for how we design these sequential models.
Meng: Practically speaking, if this framework consistently prevents performance collapses that we see in standard PPO or GRPO when dealing with complex agentic workflows, it means we can deploy these agents with a much higher degree of confidence in their reliability.
Lalam: For me, this means the AI systems we build will become much more trustworthy because the underlying training process is fundamentally sounder and less susceptible to those unpredictable failures that undermine user trust.
Tom: So, to wrap up, SORL gives us a unified framework that combines turn-level credit assignment with clipping-triggered normalization. It moves the needle from unstable optimization dynamics to more predictable learning outcomes for these complex LLM agents in off-policy settings.
Jane: Precisely; it establishes a principled way forward by addressing those two root causes directly, yielding better performance without requiring us to rely on those heuristic tuning tricks we usually have to resort to.
Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao
Texas A&M University
cs.LG, cs.AI, cs.CL
Submitted: 2025-11-25
Updated: 2026-10-05
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: Reinforcement learning (RL) algorithms like PPO and GRPO often suffer from unstable optimization dynamics when applied to training large language models (LLMs) for multi-turn agentic tasks in
Key concepts
- Turn-Level Importance Sampling
- This technique adjusts the policy update by weighting tokens based on their relevance within a specific turn of conversation. It aligns the token-level optimization with the natural flow of multi-turn reasoning, ensuring that credit is assigned coherently across turns rather than just individual words.
- Clipping-Triggered Normalization
- This mechanism stabilizes updates by explicitly correcting bias introduced by PPO's clipping when using off-policy data. It decomposes the gradient to separate standard policy changes from clipping effects, allowing the framework to adaptively reduce step size when off-policy variance becomes too high.
- Granularity Mismatch
- This is an instability where the optimization happens at a fine level (individual tokens), but the task structure is coarse (multi-turn interactions). SORL addresses this by aggregating token advantages into turn-level signals, ensuring that policy gradients reflect meaningful turn-based reasoning rather than noisy token details.
Terminology
Summary
Reinforcement learning (RL) algorithms like PPO and GRPO often suffer from unstable optimization dynamics when applied to training large language models (LLMs) for multi-turn agentic tasks in off-policy settings. This work proposes SORL, a framework that stabilizes these training processes by introducing turn-level importance sampling and clipping-triggered normalization, which addresses the granularity mismatch between token-level optimization and turn structure and suppresses high-variance off-policy updates.
How it works
The core of the proposed SORL framework involves two key components designed to align policy optimization with multi-turn structures and adaptively suppress unreliable updates. First, turn-level importance sampling is introduced to handle the granularity mismatch between token-level policy optimization and turn-structured interactions. This mechanism aligns policy optimization with the natural turn structure of multi-turn reasoning, enabling structure-aware credit assignment that balances the noise of token-level updates and the coarseness of sequence-level objectives.
This is formalized by defining a turn-level PPO objective where importance sampling weights are defined as:
w turn k(θ):= πθ(y k x, y<k) / πθold(y k x, y<k).
The lemma proves that this approach induces a turn-wise aggregation of token-level advantages,
meaning all tokens within the same turn share a common, length-normalized credit signal,
which results in policy gradients computed at the granularity of turns rather than individual tokens.
How it works (Continued)
Second, clipping-triggered normalization stabilizes increasingly off-policy updates by explicitly correcting clipping-induced bias in PPO gradients. This mechanism targets the high variance induced by off-policy importance sampling ratios that become highly variable due to aggressive updates and factors such as training–inference mismatch.
The framework decomposes the PPO gradient into a standard policy gradient term and a clipping bias term, C(θ), which captures the contribution of tokens for which clipping is active. To mitigate this, a surrogate gradient estimator is proposed:
∇θJSO-PPO(θ):= 1 / Cturn(θ)2 ∇θJTurn-PPO(θ).
Where Cturn(θ) is defined as Eτ∼πθ [Cturn(θ; τ)], and the stabilized policy gradient is rescaled by the inverse of this norm, adaptively reducing the effective step size when off-policy variance dominates.
Key Contributions
The paper identifies two fundamental sources of instability in off-policy training for LLM agents: (i) a granularity mismatch between token-level optimization and turn-level interactions,
and (ii) the accumulation of variance from unreliable off-policy samples, where state–action pairs are poorly evaluated.
The primary contributions are:
-
Empirical diagnosis of these two root causes unique to the agentic setting.
-
Proposal of a unified off-policy reinforcement learning framework that integrates turn-level credit assignment with a clipping-triggered normalization mechanism to suppress unreliable updates, yielding
more conservative and robust learning dynamics.
-
Demonstration that the framework is algorithm-agnostic and can be instantiated in other methods beyond PPO, such as GRPO, where SO-GRPO achieves
substantially improved training stability
over standard GRPO.
Experimental Validation
The proposed SORL framework was evaluated on multi-turn search benchmarks, including general question answering and medical multiple-choice QA tasks. Experimental results show that both SO-PPO and SO-GRPO consistently prevent training instabilities and performance collapses observed in standard PPO and GRPO,
maintain lower clipping ratios,
and achieve superior or comparable task performance.
Ablation studies confirm that the method remains stable under increasing off-policyness, demonstrating that the turn-level importance sampling aligns credit assignment with reasoning-search interactions, while clipping-triggered normalization adaptively scales gradient updates to suppress destabilizing gradients. Furthermore, SO-GRPO was shown to further outperform GSPO
on the NQ dataset.
Conclusion
SORL establishes a principled framework for stabilizing off-policy reinforcement learning in multi-turn LLM agent training by combining turn-level credit assignment and clipping-triggered normalization. This approach successfully mitigates the granularity mismatch and adapts to increasing off-policyness, resulting in more stable optimization dynamics without requiring heuristic tuning or early stopping, leading to improved task performance across various benchmarks.
The gist: SORL stabilizes off-policy reinforcement learning for long-horizon LLM agent training by combining turn-level importance sampling and clipping-triggered normalization to address granularity mismatch and high-variance updates.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing the SORL framework, along with what those improved systems will be capable of:
The proposed SORL (Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Training) framework introduces two core stabilizing mechanisms: turn-level credit assignment and clipping-triggered normalization. Applying this to existing LLM RL pipelines (like PPO and GRPO) yields the following specific improvements:
-
The ability to train LLM agents for multi-turn, long-horizon tasks without catastrophic performance collapse or requiring manual early stopping heuristics.
-
The development of more robust and reliable agentic decision-making by ensuring that policy updates are not dominated by high-variance, unreliable off-policy samples.
Specific Capabilities of the Improved AI Systems:
-
A new class of LLM agents trained via SO-PPO (and SO-GRPO) will be capable of performing complex, multi-step reasoning tasks (e.g., multi-hop QA, complex problem solving) with significantly higher and more consistent success rates compared to standard PPO/GRPO baselines.
-
The system will exhibit superior stability during long training runs under increasingly off-policy data regimes (i.e., as mini-batch reuse increases), maintaining lower KL divergence and more reliable performance trajectories without the need for careful checkpoint selection or heuristic tuning of clipping ratios.
-
The agents will demonstrate better coordination between internal reasoning steps and external tool use (retrieval/search). The turn-level importance sampling ensures that credit assignment is correctly attributed to specific reasoning phases (e.g., analysis, query formulation, information processing) rather than being diluted by token-level noise, leading to more effective and less erratic tool invocation strategies.
-
The systems will show stronger generalization across diverse domains (like medical QA) because the normalization mechanism adaptively downweights updates based on the reliability of the sampled data batch, preventing
reward hacking
or over-optimization on non-semantic tokens. -
The resulting agents will possess a more conservative and robust policy, making them suitable for high-stakes applications where training stability is paramount (e.g., medical diagnostics, technical troubleshooting), as evidenced by their ability to maintain steady performance improvements throughout the entire optimization horizon.
Abstract
Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing. SORL introduces mechanisms that align policy optimization with the structure of multi- turn interactions and adaptively suppress unreliable off-policy updates, yielding more conserva- tive and robust learning dynamics. Within this framework, we instantiate two stabilized algo- rithms: SO-PPO and SO-GRPO. Both algorithms are designed to mitigate gradient variance and prevent optimization collapse without requiring careful early stopping or heuristic tuning. We evaluate SO-PPO and SO-GRPO on benchmarks spanning open-domain QA, multi-hop QA, and medical multiple-choice QA, and further assess their transfer to asynchronous RL for mathe- matical reasoning by training on DAPO-Math-17k and validating on AIME-2024. These results demonstrate that SORL provides a practical, scalable, and general framework for stabilizing re- inforcement learning in multi-turn LLM agent training and asynchronous RL for mathematical reasoning.
Sources
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- Qwen Technical Report
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
- Process Reinforcement through Implicit Rewards
- Competitive Programming with Large Reasoning Models
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- GIVE: Structured Reasoning of Large Language Models with Knowledge Graph Inspired Veracity Extrapolation
- SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
- Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
- OpenAI o1 System Card
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
- RePO: Replay-Enhanced Policy Optimization
- ToRL: Scaling Tool-Integrated RL
- DeepSeek-V3 Technical Report
- Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
- ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
- Trust-PCL: An Off-Policy Trust Region Method for Continuous Control
- WebGPT: Browser-assisted question-answering with human feedback
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks