RL-finetuning LLMs from on- and off-policy data with a single algorithm

arXiv:2503.19612 · cs.LG, stat.ML · Submitted 2025-03-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "RL-finetuning LLMs from on- and off-policy data with a single algorithm".

Jane: The proposed Any-Generation Reward Optimization (AGRO) algorithm introduces a single reinforcement learning framework capable of fine-tuning large language models effectively across both on-policy and off-policy data settings by leveraging the…

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, shifting over to the title and authors of this paper, "RL-finetuning LLMs from on- and off-policy data with a single algorithm," it really captures the essence of what they're proposing.

Jane: That title tells us immediately that the main contribution is unifying fine-tuning methods that traditionally have been separate, specifically addressing how we use both on and off policy data streams simultaneously.

Lu: The authors are Yunhao Tang, Taco Cohen, David W. Zhang, and Michal Valko; they're clearly tackling a problem central to modern alignment research by proposing this single algorithm called AGRO.

Meng: They're aiming for a sample-based policy gradient approach that has theoretical guarantees on its convergence, which is something we always look for when moving beyond just empirical success.

Lalam: It’s exciting because it moves away from needing complex auxiliary value functions, suggesting a more direct path to finding the optimal policy.

The paper's summary: Tom: Now, let's talk about the actual summary of this paper; essentially, they introduce AGRO and show how it uses generation consistency as its central theoretical pillar for fine-tuning large language models.

Jane: So, in simple terms, the authors are stating that if a policy is consistent across any possible generation of the model, then it must be optimal according to their framework.

Lu: That concept of consistency leading to an optimality condition where the reward function depends only on the prompt is a really powerful theoretical insight for understanding what makes an AI model behave well.

Meng: When they derive these loss functions, L(pi) for on-policy and L mu(pi) for off-policy, it means they have concrete mathematical targets to minimize when searching for that optimal policy.

Lalam: Minimizing both of these losses globally is what leads to the final optimal policy, which simplifies the search process significantly compared to trying them separately.

The paper's improvements: Tom: Looking at the specific improvements they detail, AGRO focuses on deriving two primary loss functions based on that consistency condition—the on-policy loss and the off-policy loss.

Jane: What’s really interesting is how they show that globally minimizing both L(pi) and L mu(pi) actually results in finding the optimal policy, which is a very clean way to frame the optimization problem.

Lu: They also provide some gradient estimates for policy optimization, like Lemma three which simplifies things when we consider how to actually calculate those gradients in practice for large language models <ref:2503.19612#pg0>.

Meng: The paper then moves into token-level implementation by presenting three unbiased stochastic estimates for a regularization term g(pi), which is crucial because sequence-level generations need careful handling.

Lalam: That token-level variance reduction is what makes the algorithm practical, allowing it to handle the high dimensionality of LLMs without getting bogged down in noise during training.

Conclusion: Tom: So, to wrap up on this paper, we're seeing a unified framework that handles on-policy and off-policy data streams using generation consistency as the core idea for finding the optimal policy.

Jane: The implication is that we can potentially use whatever data we have available—whether it’s fresh human feedback or older datasets—and apply this one algorithm to get better results than existing methods.

Lu: It really opens up possibilities for applying this structure to more complex reward learning scenarios, especially when thinking about how consistency might generalize beyond just the reward function itself.

Meng: I think the practical impact is significant because it promises reduced sample complexity and improved training stability compared to older policy gradient methods we've seen.

Lalam: For me, the biggest cultural implication is that this suggests a more robust way to build aligned models that are less sensitive to the specific data regime we use for training them.

Tom: So, that’s what it is; a single algorithm tackling the complexity of on- and off-policy RL fine-tuning through the lens of generation consistency in "RL-finetuning LLMs from on- and off-policy data with a single algorithm." We're ready to hear what's next.

Yunhao Tang, Taco Cohen, David W. Zhang, Michal Valko, Remi Munos

cs.LG, stat.ML

Submitted: 2025-03-25

Updated: 2025-03-28

Importance score: 81/100

The gist: The proposed Any-Generation Reward Optimization (AGRO) algorithm introduces a single reinforcement learning framework capable of fine-tuning large language models effectively across both on-policy

Key concepts

Generation Consistency
This concept states that the optimal policy should behave consistently across any possible way a model generates text. If you generate two versions of the output from the same prompt, the optimal policy should maintain a predictable relationship between those two generations, making it easier to find the best settings.
KL Regularization
This is a mathematical constraint used during fine-tuning to keep the new model's behavior close to its original reference model. It prevents drastic changes in how the model responds while still allowing it to learn from new data, ensuring stability during training.
On-policy vs. Off-policy Loss
These are two different ways of calculating the loss function used to train the model. On-policy loss uses data generated by the current policy, while off-policy loss uses data generated by a different policy. AGRO minimizes both to find the best overall model.
Generation Consistency Condition
This is a theoretical finding showing that at the optimal point, a specific mathematical relationship involving rewards and log probabilities remains constant regardless of which specific generation you test. This simplifies finding the best policy because it means the reward calculation depends only on the prompt itself.

Terminology

Summary

The proposed Any-Generation Reward Optimization (AGRO) algorithm introduces a single reinforcement learning framework capable of fine-tuning large language models effectively across both on-policy and off-policy data settings by leveraging the concept of generation consistency. This method is significant because it derives optimal policies through sample-based learning with theoretical guarantees, demonstrating improved performance on mathematical reasoning datasets over existing baseline algorithms.

The gist

AGRO leverages the concept of generation consistency, which states that the optimal policy satisfies a notion of consistency across any possible generation of the model.

Reinforcement Learning Formulation and Consistency Condition

The paper frames language model fine-tuning as a regularized policy optimization problem: maximizing the objective function defined by Equation (1), which aims to maximize expected reward while maintaining proximity to a reference policy through KL regularization. A central theoretical finding is Theorem 1, which establishes generation consistency at optimality: for any generation y, the quantity r(x, y) − β log π∗(yx)πref(yx) does not depend on y; it is a function of the prompt x only. This implies that the optimal policy satisfies a specific relationship where the regularized reward depends only on the prompt.

Derivation of Loss Functions

Based on generation consistency, two primary loss functions are introduced to search for the optimal policy:

  1. The on-policy loss, defined as:

L(π) = 1/2 Ex∼ρ [Vy∼π(·x) Rπβ(x, y)].

  1. The off-policy loss, defined as:

Lµ(π) = 1/2 Ex∼ρ [Vy∼µ(·x) Rπβ(x, y)].

The paper demonstrates that globally minimizing both L(π) and Lµ(π) leads to the optimal policy.

Gradient Estimates for Policy Optimization

The gradient of the off-policy loss, ∇Lµ(π), is given by Lemma 3:

∇Lµ(π) = −βE [Rπβ(x, y)] − Rπ(x, µ) / ∇ log π(yx).

When the sampling is on-policy (µ = π), this gradient simplifies to an expression equivalent to the RLHF gradient from Equation (1), aligning with the standard on-policy RL objective. The full on-policy AGRO gradient, ∇L(π), is decomposed into two terms: ∇PDL(π) and ∇LRL(π).

Token-Level Implementation and Variance Reduction

For LLM applications, the sequence-level generations are adapted to a token level. Three unbiased stochastic estimates for the regularization term g(π) are presented:

  1. gˆ1: The vanilla estimate of g(π).

  2. gˆ2: An estimate that ignores lower-triangular terms of gˆ1, which are expected to be zero in expectation, reducing variance compared to gˆ1.

  3. gˆ3: An estimate that further reduces variance by computing the conditional expectation over each single next token, utilizing the available information from the forward pass.

Experimental Results and Performance

Experiments on the mathematical reasoning dataset MATH show that AGRO outperforms baseline algorithms in both on-policy and off-policy settings. Specifically, Figure 3 indicates that on-policy ARGO achieves the best performance overall, with a +7% lift in performance compared to πref, while off-policy AGRO seems to obtain better performance overall compared to KL-regularized policy gradient algorithm. The analysis shows that while on-policy algorithms drive up the KL divergence more, they achieve faster progress toward training reward. Strong regularization (large β) makes the algorithm more KL-efficient but can lead to instability during off-policy learning.

Theoretical Guarantees

The paper provides theoretical guarantees through Lemma 5, which states that the KL divergence is related to performance under the regularized policy as: βKL(π, π∗) = G(π∗) − G(π). This connects the policy optimization objective directly to the KL divergence between policies. Furthermore, Proposition 4 confirms that for off-policy AGRO with on-policy sampling, ∇Lb(π) is an unbiased estimate of ∇L(π), ensuring it leads to the optimal policy π∗.

Extensions and Future Directions

The formulation can be extended to pairwise preference data using the Bradley-Terry assumption, which implies that reward differences are related to the log ratio of policies. The generation consistency condition also relates to general consistency for regularized value-based learning. Potential future work includes a more careful investigation into "the source of off-policy instability and how one might leverage complementary techniques such as importance sampling.

Improvements for AI systems

As a fastidious researcher, I have analyzed the AGRO framework presented in this paper. The core contribution is establishing generation consistency as an optimality condition for RLHF and deriving a unified algorithm (AGRO) that effectively leverages both on-policy and off-policy data streams without requiring complex auxiliary value functions.

Here are the specific improvements to AI systems that can be made using this paper, categorized by capability:


)

Large Language Model (LLM) Fine-Tuning for Mathematical Reasoning and Complex Instruction Following.

  1. A unified fine-tuning pipeline capable of optimizing LLMs using both on-policy human feedback (RLHF/DPO style) and static, pre-existing datasets (like MATH or other reasoning corpora).

  2. The ability to adapt the learning algorithm seamlessly between these two data regimes without switching codebases or fundamentally altering the objective function structure.

Specific improvements achievable:

  1. Inference on complex mathematical proofs and problem-solving tasks where high accuracy is required, leveraging the improved performance observed on the MATH dataset (+7% lift in evaluation performance).

  2. Instruction following for multi-step, chain-of-thought reasoning tasks, as the reward function directly optimizes for correct reasoning paths rather than just final answers.

  3. Handling complex preference data (Pairwise Preference) by directly optimizing the policy based on the Bradley-Terry model implied by generation consistency, leading to more robust alignment with human preferences.

)

Improved Sample Efficiency and Training Stability in RLHF/Alignment.

  1. Significantly reduced sample complexity during fine-tuning, as AGRO is demonstrated to be more data-efficient than vanilla KL-regularized policy gradient, requiring fewer total training steps (epochs).

  2. Enhanced training stability, particularly in off-policy settings where the standard KL regularization approach often leads to performance crashes; AGRO maintains consistent performance across different learning dynamics.

  3. Training high-quality alignment models with smaller datasets or fewer total gradient updates, drastically reducing computational costs associated with expensive LLM fine-tuning runs.

  4. Deployment of RLHF agents in production environments where data collection is costly, as the algorithm can effectively utilize existing offline interaction data (off-policy) while converging to the optimal policy faster than traditional methods.


)

Robust and Low-Variance Gradient Estimation for Large Language Models.

  1. Implementation of highly stable gradient estimators (like those derived from generation consistency) that minimize variance without introducing significant bias, especially when dealing with high-dimensional sequence data (token level).

  2. The ability to implement low-variance estimates for complex regularization terms in the token-level policy gradient, which is crucial for scaling AGRO to very long sequences and high-resolution token predictions.

  3. Training LLMs on extremely long contexts or sequence lengths (e.g., 4K+ tokens) where standard gradient estimation suffers from high variance, leading to more accurate token-level policy updates.

  4. More precise control over the trade-off between reward maximization and KL divergence minimization during training, resulting in policies that are both highly aligned with human preferences and maintain strong generative capabilities.


)

Generalization Across On-Policy and Off-Policy Data Regimes.

  1. Creation of a single, general-purpose RL fine-tuning algorithm (AGRO) that is inherently agnostic to whether the data used for policy updates is generated by the current model (on-policy) or an external behavior policy (off-policy).

  2. The ability to derive optimal policies by leveraging diverse data sources simultaneously—using on-policy experience for immediate feedback and off-policy data for broad coverage—without needing separate specialized algorithms for each scenario.

  3. Developing a single, robust fine-tuning tool that can be used in various industrial settings (e.g., fine-tuning on a small high-quality human preference set on-policy, then leveraging a large, older dataset off-policy).

  4. Mitigating the online vs. offline trade-off by ensuring that the theoretical convergence guarantees hold for both scenarios under the AGRO framework, allowing practitioners to choose data sources based purely on cost/availability rather than algorithmic limitations.

Sources

Related papers