Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
summary
The gist
Reinforcement Learning's efficacy in refining robotic manipulation policies is currently bottlenecked by the difficulty of designing generalizable, dense reward functions, which this paper addresses
In short
The paper addresses difficulty in designing good reward functions for training robots using Reinforcement Learning. It proposes Large Reward Models (LRMs) that adapt vision-language models to create online reward generators. These LRMs produce a complex, multi-faceted reward signal based on visual observations, improving policy refinement through process, progress, and completion feedback.
Key concepts
- Large Reward Models (LRMs)
- These are specialized versions of foundation Vision-Language Models adapted for generating rewards. They are trained to map visual inputs and task descriptions directly into a numerical reward signal. The paper uses them to create three distinct reward types: temporal contrastive, absolute progress, and task completion.
- Temporal Contrastive Reward (rcont)
- This reward measures relative progress by comparing two consecutive frames in time. It determines if the current state is closer to the goal than the previous state. This method provides dense feedback on movement direction without needing an exact absolute score, helping robots understand which way to move relative to completion.
- Online Policy Refinement
- This is the process of using Reinforcement Learning (RL) to improve a robot's behavior while it is actively interacting with the environment. The LRMs act as an online engine, providing immediate reward signals based on what the robot sees, allowing the policy to learn and adjust its actions iteratively.
Terminology used across episodes
This episode discusses
- Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models · Paper Radio
- RoboReward: General-Purpose Vision-Language Reward Models for Robotics
- Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
- ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
- Robot Learning from a Physical World Model
- Seeing the Wind from a Falling Leaf
- Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
- A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning
- RLIF: Interactive Imitation Learning as Reinforcement Learning
- The Ingredients of Real-World Robotic Reinforcement Learning
- Continuously Improving Mobile Manipulation with Autonomous Real-World RL
- Octo: An Open-Source Generalist Robot Policy
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- pi* 0.6: a VLA That Learns From Experience
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
- Eureka: Human-Level Reward Design via Coding Large Language Models
- TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
The paper
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models · Read on arXiv
USC Physical Superintelligence Lab 2 Toyota Research Institute
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Large Reward Models".
Dev: Reinforcement Learning's efficacy in refining robotic manipulation policies is currently bottlenecked by the difficulty of designing generalizable, dense reward functions,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: I was really interested in the title of this paper, "Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models," because it sounds like it tackles that big problem we always run into with making robots learn to do complex tasks. It suggests they're using these large models to create rewards online rather than just having a fixed set of rules.
Dev: I agree, Rosa, the focus on generalizable reward generation is exactly where things get stuck in policy refinement; if the reward isn't robust, the policy won't learn reliably. The VLM aspect makes it sound like they are using those powerful language models to bridge that gap between high-level instructions and low-level visual feedback.
Taro: From an autonomy researcher's view, I wonder how generalizable that really means in practice; does it mean a policy trained for one type of task can immediately apply this reward system to a completely different physical setup?
Rosa: That’s exactly my question, Taro; I need to know if this works outside of the highly controlled lab environment. If it needs constant retraining for every new physical setup, then its utility in real-world field robotics is limited.
Dev: The paper mentions they trained the backbone on a large-scale dataset covering real-robot trajectories and diverse simulated environments, which suggests they aimed for some level of transferability across domains.
Taro: If the training data includes diverse simulated environments, that might help with generalization, but I’m still skeptical about how well it handles true novelty where the visual context is entirely new to the model.
Rosa: It seems like they are aiming to move away from manual reward engineering by using this VLM framework to generate a richer signal directly from what the robot sees at any given moment.
Dev: That density of feedback sounds promising for stabilizing learning, but we have to keep an eye on how that reward stream translates into actionable control signals without introducing latency issues in the loop rate.
The paper's summary: Rosa: So, what they’re actually proposing is a framework where they take a foundation Vision-Language Model and adapt it to act as an online reward generator for refining robot policies. It’s not just one simple reward; they create three distinct types of signals from the visual observations.
Dev: I see that they are decomposing the evaluation into process, completion, and temporal contrastive rewards; that structured signal decomposition is smart because it gives us different kinds of feedback at different times during interaction.
Taro: The concept of a temporal contrastive reward sounds particularly interesting for evaluating relative progress; it suggests the AI can judge which state is closer to the goal without relying on an absolute score, which addresses some calibration issues.
Rosa: Exactly, Taro; that relative ranking approach, alongside a regression task for absolute progress and a binary check for completion, gives the policy very detailed information about its performance at every step.
Dev: And from an engineering standpoint, having three different reward modalities means we have multiple ways to supervise the policy during training; we can tune how much weight each reward contributes to the overall objective function.
Taro: I’m curious about how this structured feedback handles situations where things go wrong unexpectedly; if the world misbehaves, does that structured signal help the AI recover faster than a standard sparse reward?
Rosa: Well, they suggest that by anchoring policy updates in these semantically grounded rewards—based on visual cues like object displacement or proximity—the system can resolve sub-optimal behaviors much more effectively.
Dev: That sounds like it tackles the credit assignment problem head-on; instead of just knowing the final result was bad, the policy gets feedback on *why* it got there based on its visual perception.
The paper's improvements: Rosa: The paper highlights several key improvements, starting with how they specialize a foundation model like Qwen3-VL-8B-Instruct using LoRA to create these three specific reward modalities. That fine-tuning process is crucial for making the VLM actually perform the intended functions.
Dev: I noticed they describe training three different specialization paths: a Contrastive Discrimination Model, a Progress Estimation Model, and a Completion Judgment Model; that layered approach shows they didn't just try to use one monolithic reward function.
Taro: The idea of using Direct Preference Optimization for the contrastive reward, tying it to verifiable physical interactions like object displacements, seems like a solid way to ensure the feedback is grounded in reality rather than just language semantics.
Rosa: That’s right; they want to make sure that when the model learns what "progress" means, it’s explicitly linked to measurable physical movement between frames, which really strengthens the connection to physical reality.
Dev: And for those specific models, they use Supervised Fine-Tuning to maximize likelihood for reasoning and prediction tasks; that suggests a careful process of training each component separately before integrating them into the online loop.
Taro: I’m looking forward to seeing how robust this is when we test it on unseen environments, because their goal is zero-shot generalization across diverse physical settings, which is a big hurdle in autonomy.
Rosa: That zero-shot claim is ambitious; they state that by bridging high-level semantic instructions with fine-grained visual cues across human and robotic domains, these perception capabilities form the foundation for this generalization.
Conclusion: Dev: So, to wrap up this discussion on "Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models," the core idea is that adapting VLMs into online reward generators provides a robust way to guide policy refinement using process, completion, and temporal contrastive rewards.
Rosa: That's right; the framework successfully moves beyond just post-hoc trajectory evaluation by formulating a multifaceted reward signal directly from visual observations during active interaction. It gives the robot continuous guidance based on what it perceives moment by moment.
Taro: I think this approach is significant because it addresses how we give agents dense feedback in complex, long-horizon tasks where simple binary success or failure isn't enough information for learning to occur efficiently.
Dev: From an engineering perspective, the refinement process using Proximal Policy Optimization and GAE, anchored by these LRM-generated rewards, allows the policy to update itself based on semantically grounded feedback in real-time during execution.
Rosa: Indeed, this system has the potential to allow robots to achieve high-precision manipulation autonomously from an imitation learning baseline in a relatively small number of reinforcement learning iterations.
Taro: If we can successfully deploy this for field robotics, it means we could have agents that are far more adaptive when they encounter unexpected physical disturbances or novel objects outside of their training set.
Dev: We still need to focus on the computational aspect, specifically ensuring the interval-hold strategy and the K steps for querying the LRM keep up with our required loop rates without introducing unacceptable latency.
Rosa: That’s a fair point, Dev; while the reward generation is powerful, its real-world viability hinges on that online inference speed and reliability under real operational conditions.
Taro: It's exciting to see how this structured feedback mechanism helps resolve those credit assignment issues in RL, anchoring the policy updates in verifiable physical feedback from the LRM’s reasoning about displacement and progress.
Dev: So, while we see significant gains in metrics like increasing Kendall’s tau by fifteen point three percent for contrastive rewards and dropping MAE by twenty percent for progress estimation on benchmarks like ManiSkill3, the paper also notes that the LRM's effectiveness is tied directly to the quality and diversity of its training data, which is a necessary caveat.
Rosa: Absolutely; the paper states that these gains are achieved because they leveraged a large-scale, multi-source dataset encompassing real-world robot trajectories and human-object interactions. This shows the dependence on rich data for achieving those results in the Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models paper.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications