Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models

arXiv:2603.16065 · cs.RO, cs.AI · Submitted 2026-03-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Large Reward Models".

Dev: Reinforcement Learning's efficacy in refining robotic manipulation policies is currently bottlenecked by the difficulty of designing generalizable, dense reward functions,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: I was really interested in the title of this paper, "Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models," because it sounds like it tackles that big problem we always run into with making robots learn to do complex tasks. It suggests they're using these large models to create rewards online rather than just having a fixed set of rules.

Dev: I agree, Rosa, the focus on generalizable reward generation is exactly where things get stuck in policy refinement; if the reward isn't robust, the policy won't learn reliably. The VLM aspect makes it sound like they are using those powerful language models to bridge that gap between high-level instructions and low-level visual feedback.

Taro: From an autonomy researcher's view, I wonder how generalizable that really means in practice; does it mean a policy trained for one type of task can immediately apply this reward system to a completely different physical setup?

Rosa: That’s exactly my question, Taro; I need to know if this works outside of the highly controlled lab environment. If it needs constant retraining for every new physical setup, then its utility in real-world field robotics is limited.

Dev: The paper mentions they trained the backbone on a large-scale dataset covering real-robot trajectories and diverse simulated environments, which suggests they aimed for some level of transferability across domains.

Taro: If the training data includes diverse simulated environments, that might help with generalization, but I’m still skeptical about how well it handles true novelty where the visual context is entirely new to the model.

Rosa: It seems like they are aiming to move away from manual reward engineering by using this VLM framework to generate a richer signal directly from what the robot sees at any given moment.

Dev: That density of feedback sounds promising for stabilizing learning, but we have to keep an eye on how that reward stream translates into actionable control signals without introducing latency issues in the loop rate.

The paper's summary: Rosa: So, what they’re actually proposing is a framework where they take a foundation Vision-Language Model and adapt it to act as an online reward generator for refining robot policies. It’s not just one simple reward; they create three distinct types of signals from the visual observations.

Dev: I see that they are decomposing the evaluation into process, completion, and temporal contrastive rewards; that structured signal decomposition is smart because it gives us different kinds of feedback at different times during interaction.

Taro: The concept of a temporal contrastive reward sounds particularly interesting for evaluating relative progress; it suggests the AI can judge which state is closer to the goal without relying on an absolute score, which addresses some calibration issues.

Rosa: Exactly, Taro; that relative ranking approach, alongside a regression task for absolute progress and a binary check for completion, gives the policy very detailed information about its performance at every step.

Dev: And from an engineering standpoint, having three different reward modalities means we have multiple ways to supervise the policy during training; we can tune how much weight each reward contributes to the overall objective function.

Taro: I’m curious about how this structured feedback handles situations where things go wrong unexpectedly; if the world misbehaves, does that structured signal help the AI recover faster than a standard sparse reward?

Rosa: Well, they suggest that by anchoring policy updates in these semantically grounded rewards—based on visual cues like object displacement or proximity—the system can resolve sub-optimal behaviors much more effectively.

Dev: That sounds like it tackles the credit assignment problem head-on; instead of just knowing the final result was bad, the policy gets feedback on *why* it got there based on its visual perception.

The paper's improvements: Rosa: The paper highlights several key improvements, starting with how they specialize a foundation model like Qwen3-VL-8B-Instruct using LoRA to create these three specific reward modalities. That fine-tuning process is crucial for making the VLM actually perform the intended functions.

Dev: I noticed they describe training three different specialization paths: a Contrastive Discrimination Model, a Progress Estimation Model, and a Completion Judgment Model; that layered approach shows they didn't just try to use one monolithic reward function.

Taro: The idea of using Direct Preference Optimization for the contrastive reward, tying it to verifiable physical interactions like object displacements, seems like a solid way to ensure the feedback is grounded in reality rather than just language semantics.

Rosa: That’s right; they want to make sure that when the model learns what "progress" means, it’s explicitly linked to measurable physical movement between frames, which really strengthens the connection to physical reality.

Dev: And for those specific models, they use Supervised Fine-Tuning to maximize likelihood for reasoning and prediction tasks; that suggests a careful process of training each component separately before integrating them into the online loop.

Taro: I’m looking forward to seeing how robust this is when we test it on unseen environments, because their goal is zero-shot generalization across diverse physical settings, which is a big hurdle in autonomy.

Rosa: That zero-shot claim is ambitious; they state that by bridging high-level semantic instructions with fine-grained visual cues across human and robotic domains, these perception capabilities form the foundation for this generalization.

Conclusion: Dev: So, to wrap up this discussion on "Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models," the core idea is that adapting VLMs into online reward generators provides a robust way to guide policy refinement using process, completion, and temporal contrastive rewards.

Rosa: That's right; the framework successfully moves beyond just post-hoc trajectory evaluation by formulating a multifaceted reward signal directly from visual observations during active interaction. It gives the robot continuous guidance based on what it perceives moment by moment.

Taro: I think this approach is significant because it addresses how we give agents dense feedback in complex, long-horizon tasks where simple binary success or failure isn't enough information for learning to occur efficiently.

Dev: From an engineering perspective, the refinement process using Proximal Policy Optimization and GAE, anchored by these LRM-generated rewards, allows the policy to update itself based on semantically grounded feedback in real-time during execution.

Rosa: Indeed, this system has the potential to allow robots to achieve high-precision manipulation autonomously from an imitation learning baseline in a relatively small number of reinforcement learning iterations.

Taro: If we can successfully deploy this for field robotics, it means we could have agents that are far more adaptive when they encounter unexpected physical disturbances or novel objects outside of their training set.

Dev: We still need to focus on the computational aspect, specifically ensuring the interval-hold strategy and the K steps for querying the LRM keep up with our required loop rates without introducing unacceptable latency.

Rosa: That’s a fair point, Dev; while the reward generation is powerful, its real-world viability hinges on that online inference speed and reliability under real operational conditions.

Taro: It's exciting to see how this structured feedback mechanism helps resolve those credit assignment issues in RL, anchoring the policy updates in verifiable physical feedback from the LRM’s reasoning about displacement and progress.

Dev: So, while we see significant gains in metrics like increasing Kendall’s tau by fifteen point three percent for contrastive rewards and dropping MAE by twenty percent for progress estimation on benchmarks like ManiSkill3, the paper also notes that the LRM's effectiveness is tied directly to the quality and diversity of its training data, which is a necessary caveat.

Rosa: Absolutely; the paper states that these gains are achieved because they leveraged a large-scale, multi-source dataset encompassing real-world robot trajectories and human-object interactions. This shows the dependence on rich data for achieving those results in the Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models paper.

USC Physical Superintelligence Lab 2 Toyota Research Institute

cs.RO, cs.AI

Submitted: 2026-03-17

Updated: 2026-09-30

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 80/100

The gist: Reinforcement Learning's efficacy in refining robotic manipulation policies is currently bottlenecked by the difficulty of designing generalizable, dense reward functions, which this paper addresses

Key concepts

Large Reward Models (LRMs)
These are specialized versions of foundation Vision-Language Models adapted for generating rewards. They are trained to map visual inputs and task descriptions directly into a numerical reward signal. The paper uses them to create three distinct reward types: temporal contrastive, absolute progress, and task completion.
Temporal Contrastive Reward (rcont)
This reward measures relative progress by comparing two consecutive frames in time. It determines if the current state is closer to the goal than the previous state. This method provides dense feedback on movement direction without needing an exact absolute score, helping robots understand which way to move relative to completion.
Online Policy Refinement
This is the process of using Reinforcement Learning (RL) to improve a robot's behavior while it is actively interacting with the environment. The LRMs act as an online engine, providing immediate reward signals based on what the robot sees, allowing the policy to learn and adjust its actions iteratively.

Terminology

Summary

Reinforcement Learning's efficacy in refining robotic manipulation policies is currently bottlenecked by the difficulty of designing generalizable, dense reward functions, which this paper addresses by proposing Large Reward Models (LRMs) that adapt foundation Vision-Language Models (VLMs) into online reward generators.

The gist: We propose a framework for online policy refinement by adapting foundation VLMs into online reward generators to formulate a multifaceted reward signal comprising process, completion, and temporal contrastive rewards based on current visual observations.

Framework Overview

The proposed framework utilizes Reinforcement Learning (RL) to refine a robotic policy πϕ by maximizing the expected return J(πϕ) = Eτ∼πϕ "X T t=0 γ t r(It, d), where r is a generalizable reward signal mapped directly from visual observations I. This mapping is achieved by constructing a structured dataset D =

D =

for training specialized Large Reward Models (LRMs) across three functional modalities: Temporal Contrastive (rcont), Absolute Progress (rprog), and Task Completion (rcomp). This establishes a direct forward mapping rm = LRMm(It, d), which explicitly links the visual input I, the semantic task d, and the resulting reward r through the LRM’s internal reasoning.

Reward Signal Formulation

The framework designs a tri-faceted reward structure to decompose task evaluation into three complementary formats:

  1. Temporal Contrastive Reward (rcont): Designed for relative progress evaluation by comparing a pair of temporal frames to determine which state is closer to task completion, providing highly robust, dense feedback while mitigating the calibration issues of absolute scoring. This is formulated as: rcont = +1.0, if It is closer to goal than It−∆t; -1.0, if It−∆t is closer to goal than It; and 0.0 otherwise.

  2. Absolute Progress Reward (rprog): Formulated as a numerical regression task that estimates the task completion percentage based on a single input frame I, serving as: rprog ∈ 0.0, 0.1, 0.2,..., 1.0 for continuous spatial-temporal grounding.

  3. Task Completion Reward (rcomp): Provides a definitive binary assessment of the current state It by outputting a binary success signal: rcomp = (1, if semantic requirements are met; 0, otherwise).

Training Large Reward Models (LRMs)

The foundation Qwen3-VL-8B-Instruct model is specialized into LRMs via Low-Rank Adaptation (LoRA) on the structured dataset D. The training involves three distinct specialization paths:

  1. Contrastive Discrimination Model: Trained as a preference judge using Direct Preference Optimization (DPO) to align internal preferences with temporal progression, ensuring the reward rcont is derived from verifiable physical interactions—such as object displacements.

  2. Progress Estimation Model: Trained via Supervised Fine-Tuning (SFT) to maximize the likelihood of a joint reasoning-and-label sequence, compelling the model to articulate physical cues before concluding the reward rprog.

  3. Completion Judgment Model: Optimized via SFT to directly predict the binary success signal rcomp, preserving its ability to verify goal satisfaction.

Online Policy Refinement

The frozen LRMs serve as an online reward engine to refine a base policy πϕ, initialized from an Imitation Learning (IL) baseline. The process involves:

  1. Online Interaction and Reward Integration: The policy samples actions based on the current visual observation It and task description d. To bridge the computational gap, an Interval-Hold strategy is used where LRMs are queried every K environment steps to perform the forward mapping rm = LRMm(I, d), resulting in a cached reward rt = wmrm.

  2. Policy Optimization and Refinement: The Proximal Policy Optimization (PPO) framework is used to update policy parameters ϕ by maximizing the expected return J(πϕ). The optimization objective relies on Generalized Advantage Estimation (GAE) to compute the advantage Aˆt, where δt is the temporal difference (TD) error calculated using the LRM-generated reward rt. This mechanism allows the model to resolve sub-optimal behaviors by anchoring policy updates in semantically-grounded rewards.

Evaluation and Results

The framework was evaluated on challenging long-horizon manipulation benchmarks like ManiSkill3 [9] in a purely zero-shot manner for the reward models. Evaluation showed significant gains:

(A)

The LRM significantly improves ranking correlation for the Temporal Contrastive Reward, increasing Kendall’s τ and Spearman’s ρ by 15.3%. For the Absolute Progress Reward, Mean Absolute Error (MAE) drops by 20.0% and RMSE decreases by 19.

Improvements for AI systems

Here are specific improvements to AI systems based on the proposed Large Reward Models (LRMs) framework:

  1. A foundation model-based robot policy refinement system capable of achieving high-precision, long-horizon manipulation tasks autonomously from a pre-trained Imitation Learning (IL) baseline in just 30 Reinforcement Learning (RL) iterations.

  2. An AI system that generates dense, frame-level reward signals by adapting foundation Vision Language Models (VLMs) into three specialized modalities: Temporal Contrastive Reward, Absolute Progress Reward, and Task Completion Reward.

  3. A robotic control agent that utilizes a closed-loop refinement process where the policy is updated via Proximal Policy Optimization (PPO) using LRM-generated rewards to correct sub-optimal behaviors in real-time during physical execution.

  4. An AI system capable of performing zero-shot generalization across diverse, unseen physical environments by leveraging a multi-domain dataset (real robot trajectories, human interaction data, and simulated benchmarks) for LRM training.

  5. A policy refinement mechanism that resolves the credit assignment problem in RL by anchoring policy updates in semantically grounded visual feedback derived from the LRM's reasoning about object displacement and task progress.

  6. A system that can autonomously filter successful trajectories from real-world hardware rollouts using the Task Completion Reward as an automated sparse reward classifier, significantly improving sample efficiency during deployment refinement.

Sources

Related papers