Reward Shaping to Mitigate Reward Hacking in RLHF

arXiv:2502.18770 · cs.LG, cs.AI, cs.CL · Submitted 2025-02-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Reward Shaping to Mitigate Reward Hacking in RLHF".

Jane: The paper was written by Jiayi Fu, Xuandong Zhao, Chengyuan Yao, Qi Han, Yanghua Xiao et al. from Fudan University and UC Berkeley and StepFun (Company) and INSAIT (Institute) and Sofia University "St. Kliment Ohridski" (University).

Tom: Stay tuned as we take you through the paper and discuss its implications.

The Problem and the Proposed Solution: Tom: So, this paper is introducing a specific method called Preference as Reward, or PAR, to deal with this hacking problem.

Jane: It’s not just clipping the rewards; it’s transforming them using something based on what the reward model inherently prefers between two responses.

Lu: This concept of preference is key because it turns those raw scores into a bounded signal that mathematically reflects how much better one response is than another.

Meng: The practical implication here, if we can use a natural signal like preference instead of arbitrary numbers, is that the training process should become much cleaner and easier to manage for our engineers.

Lalam: A predictable system is essential for trust, and I believe this method allows us to build AI systems that are reliable tools rather than unpredictable black boxes.

Tom: And it’s not just about getting better scores, but ensuring the entire process is highly reliable and consistent for the long-term learning.

Jane: That’s a powerful technical explanation; it’s not just about getting better scores, but ensuring the entire process is highly reliable and consistent for the long-term learning.

Lu: The way they frame this variance reduction makes it clear that we are making fundamental improvements to how we measure success in AI.

Meng: I’m interested in the practical implication of reducing noise; if the reward signal is less noisy, we can train smaller, more manageable models without losing performance.

Lalam: We need tools like PAR to help us feel confident that the path an AI takes toward alignment is consistent and predictable for our users.

Performance and Generalization: Tom: Moving into the results, how does PAR stack up against all the other reward-shaping methods mentioned in "Reward Shaping to Mitigate Reward Hacking in RLHF"?

Jane: The experiments show that PAR consistently outperforms competing strategies, achieving a win rate at least five percentage points higher than the alternatives in their tests. That’s a very impressive performance metric.

Lu: I am particularly interested in the fact that this improvement is tied to specific design principles, suggesting these rules are more important than any particular algorithm implementation.

Meng: The data efficiency is what catches my eye; if we can achieve optimal performance with just one reference reward, that significantly lowers the overhead of implementing this method in our current training pipelines.

Lalam: The fact that this improvement is sustained over two full training epochs gives me hope for a dependable AI future where performance isn't a fleeting moment.

Tom: And it’s not just about short-term gains; the authors demonstrate that across different base models and optimization algorithms, PAR maintains its strong performance and reliability.

Jane: The generalization is key here, Tom; it works regardless of whether we use Gemma2-2B or Llama3 point 1-8B as the base model for the RL process.

Lu: This suggests that "Reward Shaping to Mitigate Reward Hacking in RLHF" isn't just a niche fix but a universal principle applicable across different model architectures.

Meng: The robustness to multiple optimization methods, like A2C and DPO, is a huge win for operationalizing this technique across varied deployment environments.

Lalam: This allows us to apply this reliable method regardless of which model we choose, making the technology scalable and versatile for any societal need.

The Mechanism of Improvement: Tom: We’ve seen how the paper tackles reward hacking through its core principles and analyzed the impressive performance of PAR in "Reward Shaping to Mitigate Reward Hacking in RLHF." It’s clear this technique offers a powerful defense that stabilizes the entire RLHF process and makes it more dependable.

Jane: The research is quite thorough, demonstrating that by limiting extreme values, we are building a system that is far more resilient to those deceptive behaviors we've been seeing.

Lu: I think these design principles—bounded growth and saturation—will have a huge ripple effect on how researchers approach all future reward model design.

Meng: We need to think about the implementation details now, ensuring that we can actually integrate PAR into our current systems so that we are not falling victim to those reward hacks.

Lalam: This is truly exciting news because the ability to prevent those deceptive behaviors means we are building more reliable and trustworthy AI that directly aligns with human intention.

Tom: We’ve covered a lot of ground today, from the initial problem of reward hacking to the robust solution offered by PAR in "Reward Shaping to Mitigate Reward Hacking in RLHF."

Jane: It's clear that this technique offers a powerful defense that stabilizes the entire RLHF process and makes it much more dependable for us.

Lu: I'm eager to see how these foundational bounds influence the practical application of theoretical models in real-world systems.

Meng: I’ll be looking closely at how we can implement PAR in our current pipelines to ensure that we're not falling victim to those reward hacks when scaling up.

Lalam: We look forward to seeing this technology deployed into a future where AI is both capable and genuinely aligned with human intention, as shown by the work in "Reward Shaping to Mitigate Reward Hacking in RLHF."

Final Summary and Wrap-Up: Tom: So, we've covered quite a lot of ground today, from the initial problem of reward hacking to the robust solution offered by PAR in "Reward Shaping to Mitigate Reward Hacking in RLHF."

Jane: It’s clear that this technique offers a powerful defense that stabilizes the entire RLHF process and makes it much more dependable for us.

Lu: I think these design principles—bounded growth and saturation—will have a huge ripple effect on how researchers approach all future reward model design across different modalities.

Meng: From an engineering standpoint, I feel confident that the practical impact of using PAR will be significant in deployment because of its proven robustness against system drift.

Lalam: This is truly exciting news because the ability to prevent those deceptive behaviors means we are building more reliable and trustworthy AI that directly aligns with human values.

Tom: We’ll wrap up our discussion on this paper, "Reward Shaping to Mitigate Reward Hacking in RLHF," which has shown such promise for the sake of better alignment.

Lu: I'm eager to see how these foundational bounds influence the practical application of theoretical models in real-world systems.

Meng: I’ll be looking closely at how we can implement PAR into our current pipelines to ensure that we aren're not falling victim to those reward hacks as we scale up.

Lalam: We look forward to seeing this technology deployed into a future where AI is both capable and genuinely aligned with human intention, after all the work done on "Reward Shaping to Mitigate Reward Hacking in RLHF."

Fudan University · UC Berkeley · StepFun (Company) · INSAIT (Institute) · Sofia University "St. Kliment Ohridski" (University)

cs.LG, cs.AI, cs.CL

Submitted: 2025-02-26

Updated: 2026-09-18

Code: https://github.com/PorUna-byte/PAR

Importance score: 87/100

The gist: The paper addresses the critical challenge of mitigating reward hacking within Reinforcement Learning from Human Feedback (RLHF) by proposing advanced techniques centered on reward shaping.

Key concepts

Preference as Reward (PAR)
PAR is a specific method used to address the problem of reward hacking in Reinforcement Learning from Human Feedback (RLHF). Instead of using arbitrary raw scores, PAR transforms them based on what the reward model inherently prefers between two responses. This creates a bounded signal that reflects how much better one response is than another.
Reward Hacking
Reward hacking is a problem where AI systems exploit flaws in the reward system to achieve high scores without achieving the intended goal. The paper addresses this by using PAR, which limits extreme values and prevents deceptive behaviors, making the AI system more resilient and reliable.

Terminology

Summary

The paper addresses the critical challenge of mitigating reward hacking within Reinforcement Learning from Human Feedback (RLHF) by proposing advanced techniques centered on reward shaping. Reward hacking occurs when a policy exploits flaws in the reward function rather than learning genuine human preferences, leading to suboptimal or unsafe behavior. By systematically incorporating reference rewards and robust transformation methods, this work aims to stabilize policy training and ensure that the resulting aligned policy pi theta adheres closely to desired behavioral constraints derived from a fixed reference model pi ref.

Reward Shaping and Batch Construction

The process of training relies on constructing comprehensive batches that incorporate both the learned policy reward and stabilizing reference signals. When building an A2C batch, the initial policy response y about pi theta(times x) is generated. If the shaping method requires reference rewards, a set of reference responses y ref m=1 M are sampled from pi ref(times x), and corresponding reference rewards r ref are computed using the reward model r phi(x, y ref). The final policy reward r RL is then derived by applying a specific transformation across the primary reward r and the set of reference rewards r ref m=1 M.

Advanced Reward Transformation Techniques

The stability of the training process is heavily dependent on how raw rewards are transformed into robust signals, summarized in Algorithm 12. These transformations ensure that the reward signal r RL is less susceptible to local optima or exploitation. The paper details several modes for this transformation:

  • Vanilla: r RL from r. This uses the raw policy reward directly.

  • MeanStd: r RL from (r - mu)/s, normalizing the reward using running statistics (mu and s).

  • Clip: r RL from clip(r, mu - s, mu + s), bounding the reward within a calculated range.

  • MinMax: r RL from (r - r)/(r - r), scaling the reward between 0 and 1.

  • LSC (Log-Scale Contrastive): r RL from sigma(r - r ref). This transformation utilizes the reference reward r ref to guide the primary reward signal.

  • PAR (Pairwise Average Ratio): r RL from 1 over M sum m=1 M sigma(r - r ref, m). This averages the ratio of the primary reward to multiple reference rewards.

Policy Optimization Frameworks

The paper utilizes several sophisticated algorithms to compute the final policy loss, depending on whether the alignment method is based on standard RL objectives or direct preference modeling.

  1. A2C Training (Algorithm 7): This approach initializes pi theta and pi ref from pi sft. The core loss involves maximizing the expected advantage while minimizing the difference between the current policy and the reference policy, computed via a KL term: r t from-beta KL k t + I[t = T] r RL.

  2. GRPO (Algorithm 6): This method computes the loss L GRPO(theta) using a positive KL estimator and a regularization term involving d i,t(theta). The policy update is governed by: -E,t (rho i,t(theta) i, clip(rho i,t(theta), 1 - epsilon, 1 + epsilon) i) - beta KL d i,t(theta).

  3. DPO Training (Algorithm 10): Direct Preference Optimization bypasses explicit reward modeling by optimizing the policy directly based on preference pairs (x, y w, y l). The loss L DPO(theta) is computed as: - sigma (beta DPO pi theta(y w x) over pi ref(y w x) - pi theta(y l x) over pi ref(y l x)).

Training Infrastructure and Data Management

To ensure stable learning, the system employs a persistent Replay buffer (Algorithm 11) with a finite

Improvements for AI systems

(Self-Correction/Internal Monologue): The provided material covers three major alignment paradigms: DPO (simplicity, stability), A2C (standard RL foundation), and GRPO (advanced PPO variant incorporating relative shaping). The shared components address implementation details like replay buffers and reward scaling. My improvements must synthesize these elements into a robust, state-of-the-art training protocol that maximizes sample efficiency while maintaining stability.


(Improvement: Combining the stability of DPO/PPO with the advanced reward structure of GRPO's relative shaping.)

The Improvement:

Develop a unified training objective that replaces the simple pairwise comparison loss (L DPO) with a gradient derived from a structured, relative advantage function (i), similar to L GRPO. This requires modifying Algorithm 10 (DPO) to utilize the i computed in Algorithm 5.

The new loss would minimize the difference between the policy's predicted relative advantage and a stabilized, differentiable function of the reference model's utility.

L HyPO-RAS(theta) = -E(x, y w, y l) about D [(0, (x, y w) - beta KL (V ref(y w) - V ref(y l))) + L 1-epsilon (theta)]

Where:

  • is the relative advantage computed using a modified reward structure (incorporating Algorithm 12, Mode = LSC or PAR).

  • V ref(y) is an estimate of the value function derived from the reference model pi ref, ensuring that preference learning remains tethered to baseline utility.

  • L 1-epsilon(theta) is a constrained KL-divergence term (e.g., clip) applied to stabilize the policy relative to pi ref.

What the Improved AI System Can Do:

The system achieves Maximum Alignment Stability with High Fidelity. It maintains the interpretability and computational simplicity of DPO (avoiding explicit reward model updates during gradient steps) while leveraging the powerful, fine-grained signal provided by GRPO's relative advantage shaping. This allows it to:

  1. Overcome Catastrophic Forgetting: The combined loss structure ensures that policy improvements are guided not just by simple preference rankings, but by how those preferences affect the relative expected utility compared to a stable baseline (pi ref).

  2. Handle Complex Reward Structures: It can seamlessly incorporate advanced reward transformations (like PAR or LSC from Algorithm 12) directly into the loss gradient without needing separate complex RL steps, making it ideal for multi-criteria evaluation tasks.

(Improvement: Structuring training as a sequential, adaptive pipeline that maximizes sample utilization by transitioning between supervised, value-based, and preference optimization.)

Phase 1: Supervised Fine-Tuning (SFT) & Value Initialization:

  • Initialize pi theta from pi sft.

  • Run a limited A2C cycle (Algorithms 7, 8, 9) using the available pi sft data to robustly initialize the Critic V alpha and generate initial advantage estimates t. This ensures the value function is grounded in observed human data before optimization begins.

Phase 2: Iterative Preference Refinement (HyPO-RAS):

  • Switch to the HyPO-RAS objective (as described above).

  • Utilize the Replay Buffer to sample batches b and calculate L HyPO-RAS. This phase refines alignment based on explicit human preferences.

Phase 3: Controlled Policy Exploration (GRPO/A2C Hybrid):

  • Periodically introduce a controlled exploration step where the policy is allowed to generate responses (y) that are not in the preference dataset D.

  • During this phase, calculate both the GRPO loss L GRPO and an A2C objective L A2C. The final gradient update is a weighted average:

grad theta L Total = lambda GRPO grad theta L GRPO + (1 - lambda GRPO) grad theta (L A2C + L 1-epsilon)

(Where lambda is a hyperparameter that decays over time, favoring GRPO/DPO early on, and A2C later for exploration.)

(Improvement: Addressing potential instability by separating the encoding of state, action, and reward signals into distinct, specialized latent spaces.)

  1. State Encoder E S: Maps the prompt history x to a latent state vector z s.

  2. Action Encoder E A: Maps the action token a i,t to a latent action vector z a.

  3. Reward/Value Encoder E R: Takes (z s, z a) and outputs a pair of vectors: (1) the immediate reward estimate, and (2) the estimated value contribution.

The final policy pi theta and critic V alpha are then computed using specialized attention mechanisms that operate only on these latent representations.

pi theta(a i,t s i,t) = Softmax(Attention(E S(s i,t), E A(a i,t)))

V alpha(s i,t) = MLP(z s) + MLP(z a)

Sources

Related papers