Non-Asymptotic Global Convergence of PPO-Clip
math.OC, cs.LG
Submitted: 2025-12-18
Updated: 2026-09-12
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Reinforcement learning has gained attention for modern Large Language Model post-training.
Terminology
Abstract
Reinforcement learning has gained attention for modern Large Language Model post-training. The actor-only variants of Proximal Policy Optimization (PPO) are widely applied for their efficiency. These algorithms incorporate a clipping mechanism to improve stability. Besides, a regularization term, such as the reverse KL-divergence or a more general f-divergence, is introduced to control excessive deviation from a reference policy. Despite their empirical success, a rigorous theoretical understanding of the problem and the algorithm's properties is limited. This paper advances the theoretical foundations of the PPO-Clip algorithm by analyzing a deterministic actor-only PPO algorithm within the general RL setting with f-divergence regularization under the softmax policy parameterization. We derive a non-uniform Lipschitz smoothness condition and a Łojasiewicz inequality for the considered problem. Based on these properties, we establish non-asymptotic global linear convergence in value gap for the forward KL regularizer. For the reverse KL regularizer, we derive global linear convergence from any finite softmax initialization in both value gap and squared policy distance.
Sources
- Improving alignment of dialogue agents via targeted human judgements
- A unified view of entropy-regularized Markov decision processes
- Red Teaming Language Models with Language Models
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Leverage the Average: an Analysis of KL Regularization in RL
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Related papers
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
- Incremental Learning in Mirror Flows
- Online Control via Counterfactual Tracking
- Asynchronous Replanning in Two Population Linear Quadratic Mean Field Games: Information Requirements and Stability
- Petrov-Galerkin operator inference with application to stability-encouraging identification