A Comedy of Estimators: On KL Regularization in RL Training of LLMs
cs.LG, cs.AI
Submitted: 2025-12-26
Updated: 2026-08-25
Code: https://github.com/PrimeIntellect-ai/prime-rl
License: http://creativecommons.org/licenses/by/4.0/
The gist: The reasoning performance of large language models (LLMs) can be substantially improved by training them with reinforcement learning (RL).
Terminology
Abstract
The reasoning performance of large language models (LLMs) can be substantially improved by training them with reinforcement learning (RL). The RL objective for LLM training involves a regularization term, which is the reverse Kullback-Leibler (KL) divergence between the trained policy and the reference policy. Since computing the KL divergence exactly is intractable, various estimators are used in practice to estimate it from on-policy samples. Despite its wide adoption, including in several open-source libraries, there is no systematic study analyzing the numerous ways of incorporating KL estimators in the objective and their effect on the downstream performance of RL-trained models. Recent works show that prevailing practices for incorporating KL regularization do not provide correct gradients for stated objectives, creating a discrepancy between the objective and its implementation. In this paper, we further analyze these practices and study the gradients of several estimators configurations, revealing how design choices shape gradient bias. We substantiate these findings with empirical observations by RL fine-tuning Qwen2.5-7B, Llama-3.1-8B-Instruct and Qwen3-4B-Instruct-2507 with different configurations and evaluating their performance on both in- and out-of-distribution tasks. Through our analysis, we observe that, in on-policy settings: (1) estimator configurations with biased gradients can result in training instabilities; and (2) using estimator configurations resulting in unbiased gradients leads to better performance on in-domain as well as out-of-domain tasks. We also investigate the performance resulting from different KL configurations in off-policy settings and observe that KL regularization can help stabilize off-policy RL training resulting from asynchronous setups.
Sources
- Better Estimation of the Kullback--Leibler Divergence Between Language Models
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- Process Reinforcement through Implicit Rewards
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- OpenAI o1 System Card
- Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
- Qwen2.5 Technical Report
- Proximal Policy Optimization Algorithms
- AI-Assisted Generation of Difficult Math Questions
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- On a few pitfalls in KL divergence gradient estimation for RL
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen3 Technical Report
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
- Fine-Tuning Language Models from Human Preferences
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks