Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics
cs.LG, cs.AI, cs.RO
Submitted: 2025-07-01
Updated: 2026-08-27
Comments: 25 pages, 15 figures
Code: https://github.com/rll-research/BPref
Project page: https://sunlighted.github.io/RRM-web
License: http://creativecommons.org/licenses/by/4.0/
The gist: Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments.
Terminology
Abstract
Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments. However, PbRL often suffers from poor sample efficiency, requiring extensive and costly human feedback, which limits its real-world applicability. Prior work has proposed learning a reward model from demonstrations and fine-tuning it using preferences. However, when the model is a neural network, transitioning between different loss functions across training phases often leads to unstable optimization and performance degradation. In this paper, we propose a method to effectively leverage prior knowledge with a Residual Reward Model (RRM). An RRM assumes that the true reward of the environment can be split into a sum of two parts: a prior reward and a learned reward. The prior reward is a term available before training, such as an engineering heuristic ``best guess'', a language-generated reward, or a reward function learned from inverse reinforcement learning, and the learned reward is then trained with preferences as a residual offset. Experimental results in Meta-World and DM-Control show that RRMs substantially improve the sample efficiency of common PbRL methods across various prior reward types. Furthermore, we demonstrate the practical efficacy of our method on a physical Franka Panda robot, accelerating policy learning and achieving high success rates in fewer steps than baselines.
Sources
- Residual Policy Learning
- Designing Rewards for Fast Learning
- Reward (Mis)design for Autonomous Driving
- Inverse Reward Design
- Avoiding Side Effects in Complex Environments
- Defining and Characterizing Reward Hacking
- The Ingredients of Real-World Robotic Reinforcement Learning
- B-Pref: Benchmarking Preference-Based Reinforcement Learning
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
- Reward Uncertainty for Exploration in Preference-based Reinforcement Learning
- Learning Complicated Manipulation Skills via Deterministic Policy with Limited Demonstrations
- Learning Human Objectives from Sequences of Physical Corrections
- Guiding Policies with Language via Meta-Learning
- LaND: Learning to Navigate from Disengagements
- MILE: Model-based Intervention Learning
- Learning Reward Functions by Integrating Human Demonstrations and Preferences
- Learning from Random Demonstrations: Offline Reinforcement Learning with Importance-Sampled Diffusion Models
- Loss Jump During Loss Switch in Solving PDEs with Neural Networks
- Few-Shot Preference Learning for Human-in-the-Loop RL
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks