Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

arXiv:2605.07872 · cs.CV, cs.AI · Submitted 2026-05-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Video Understanding Reward Modeling".

Jane: Multimodal reward models have advanced significantly in text and image domains,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re talking about the paper titled "Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models," and it looks like this research is really focused on solving a big problem in AI. It’s about how we teach models to judge video content better by creating a solid way to evaluate those judgments.

Jane: That sounds important, Tom; I think the title suggests they aren't just looking at making models slightly better, but rather building a whole structure for testing their ability to understand video preferences in a reliable way. It’s about establishing the right yardstick for video AI.

Lu: From a creative standpoint, I think the focus on preference judgment is really interesting because it moves beyond just telling an AI *what* happened; it forces it to learn *why* something happened in a specific context within a video.

Meng: I’m wondering what that means for practical deployment; if we have this benchmark, can we actually train models that are more reliable when they have to make choices about which video response is better?

Lalam: I see this as foundational work; by creating a structured evaluation system, we can ensure the reward signals we use to train these systems are high quality and directly map to what humans actually value in complex video scenes.

The paper's summary: Tom: Looking at the summary of "Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models," the authors basically laid out a three-part plan: they built a new benchmark called VURB, they created a massive dataset called VUP-35K, and then they trained two different reward models to work with that data.

Jane: That’s right; essentially, the paper is saying that for video understanding reward modeling to advance significantly, we need this whole unified framework—the benchmark design, the data construction pipeline, and then the training of these specialized reward models.

Lu: What struck me about their summary is how they explicitly targeted existing limitations by including long chain-of-thought reasoning traces in their VURB pairs and using a majority voting protocol for evaluation across general, long, and reasoning tasks.

Meng: So, they aren't just throwing more data at the problem; they are structuring the data collection to specifically stress-test where current reward models fall short on video understanding.

Lalam: That makes sense; by focusing on those specific areas—general understanding versus complex reasoning—they are creating a comprehensive test suite that captures the full spectrum of video comprehension needs.

The paper's improvements: Tom: Regarding the improvements they suggest, the authors focus heavily on fixing evaluation problems; they propose using VURB with its two thousand one hundred preference pairs and long chain-of-thought reasoning traces averaging one thousand one hundred forty-three tokens to see how reward models perform.

Jane: That length of reasoning is significant because it means the model has to handle much deeper inference than just looking at a final answer; it needs to follow the logic step by step.

Lu: And they also introduce the majority voting evaluation protocol specifically to mitigate position bias, which is something we’ve seen cause inconsistent results in previous evaluations of multimodal models.

Meng: From a practical standpoint, ensuring that the reward signal isn't skewed by where an answer appears on a list is crucial for building stable training loops for these reward models.

Lalam: I think the most important improvement they highlight is linking this benchmark to their data construction pipeline, VUP-35K, which they built via a fully automated process to provide large-scale supervision that was previously missing.

Conclusion: Tom: So, wrapping up the discussion on "Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models," the main implication is that we need a standardized way to evaluate video reward models that goes beyond simple metrics and includes deep reasoning traces.

Jane: I agree; this work gives us a concrete path forward by providing VURB as the new standard for testing, which should lead to more trustworthy AI systems in video understanding.

Lu: It shows how integrating robust evaluation protocols directly into the benchmark design can reveal exactly where existing reward models are lacking, which is really valuable for future research directions.

Meng: Practically speaking, this means that when we deploy these models, we can be much more confident in their performance because they’ve been trained against a rigorous set of preferences derived from this new framework.

Lalam: I think the biggest impact here is ensuring that the underlying video understanding and reasoning capabilities of the AI systems are actually enhanced, not just superficially improved by better reward scores.

South China University of Technology · Peking University · The University of Hong Kong · Tencent

cs.CV, cs.AI

Submitted: 2026-05-08

Updated: 2026-09-30

Code: https://github.com/wyclike/VURM

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 79/100

The gist: Multimodal reward models have advanced significantly in text and image domains, but progress in video understanding reward modeling remains severely limited due to a lack of robust evaluation

Key concepts

VURB
A novel benchmark specifically designed to test how well reward models judge video understanding preferences. It includes 2,100 preference pairs with long reasoning traces and uses a majority voting protocol to reduce bias in evaluation.
VUP-35K
A large dataset of 35,000 high-quality video understanding preference pairs. This data was created using an automated pipeline that samples prompts from general, long, and reasoning tasks and filters responses for quality.
VideoDRM
A discriminative reward model trained to maximize the difference between chosen and rejected responses in a ranking loss. It aims to accurately rank video understanding outputs based on preference.
VideoGRM
A generative reward model that uses Group Relative Policy Optimization (GRPO) to simultaneously produce a preference decision and an interpretable reasoning trace, optimizing for correct identification of the chosen response.

Terminology

Summary

Multimodal reward models have advanced significantly in text and image domains, but progress in video understanding reward modeling remains severely limited due to a lack of robust evaluation benchmarks and high-quality preference data. This paper addresses this critical bottleneck by proposing a unified framework encompassing benchmark design, data construction, and reward model training to establish a new standard for evaluating performance in video understanding tasks.

The Proposed Benchmark: Video Understanding Reward Bench (VURB)

The authors introduce VURB, a novel benchmark specifically designed for video understanding preference judgment. This benchmark is crucial because existing benchmarks suffer from limited scale and insufficient dimensional coverage. VURB is structured to expose key limitations of current reward models by featuring:

  1. 2,100 preference pairs with long chain-of-thought (CoT) reasoning traces, averaging 1,143 tokens.

  2. Tasks spanning general video understanding, long video understanding, and video reasoning tasks.

  3. A majority voting evaluation protocol to mitigate position bias [28].

The Data Construction Pipeline: Video Understanding Preference Dataset (VUP-35K)

To address the scarcity of high-quality preference data for video understanding, the authors construct VUP-35K, a large-scale dataset of 35K high-quality reasoning preference pairs. This construction is achieved via a fully automated and human annotator-free pipeline. The process involves:

  1. Sampling prompts from three domains: general video understanding, long video understanding, and video reasoning tasks (e.g., VideoMME, MotionBench).

  2. Generating preference responses using five thinking-enabled models (e.g., Qwen3VL-8B-Thinking), ensuring diverse CoT patterns.

  3. Applying a two-stage data quality control procedure: length consistency control (filtering pairs with relative length differences larger than 25%) and causal consistency control (removing samples where reasoning and final answers are inconsistent).

Reward Model Development: VideoDRM and VideoGRM

Building on the VUP-35K dataset, the authors train two distinct reward models: VideoDRM, a discriminative reward model, and VideoGRM, a generative reward model.

  1. For VideoDRM (discriminative), the objective is to maximize the margin between chosen and rejected responses via a ranking loss: LRM = − log sigma(rchosen − rrejected) [1].

  2. For VideoGRM (generative), they leverage Group Relative Policy Optimization (GRPO) to jointly produce a preference decision and an interpretable reasoning trace, optimizing for a binary reward (1.0 for correctly identifying the chosen response).

Evaluation and Results

The models are evaluated on VURB and VideoRewardBench. The results demonstrate that existing reward models perform surprisingly poorly on video understanding preference judgment, with most open-source models falling below 55% accuracy. The proposed VideoDRM and VideoGRM achieve state-of-the-art performance, with VideoDRM reaching an overall accuracy of 63.8% on VURB and exceeding proprietary baselines in test-time scaling experiments. Furthermore, the analysis confirms that VUP-35K not only provides core reward performance gains but also enhances the video understanding and reasoning capability of trained models.

Key Contributions

The main contributions of this work are:

**: We propose VURB, a robust benchmark tailored to video understanding preference judgment, featuring long CoT reasoning traces and majority voting evaluation. 2. **

We construct VUP-35K, a large-scale video understanding preference dataset built via an automated pipeline to address the critical scarcity of high-quality video understanding reward modeling. 3. We develop VideoDRM and VideoGRM, which achieve state-of-the-art performance on both VURB and VideoRewardBench [47] and yield significant improvements for the base model under best-of-N test-time scaling.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to AI systems by implementing the proposed framework (VURB benchmark, VUP-35K dataset, and VideoDRM/VideoGRM models):

The core improvement lies in moving beyond general multimodal understanding to achieve robust, reasoning-capable preference judgment specifically for video content.

Here are the specific improvements and capabilities:

  1. [[VURB Benchmark]]

  2. [[Long Chain-of-Thought (CoT) Reasoning Integration]]

  3. [[Majority Voting Evaluation Protocol]]


Detailed Improvements and System Capabilities:

  1. [[VURB Benchmark]]: By establishing a benchmark featuring 2,100 preference pairs with long CoT reasoning traces (averaging 1,143 tokens) across general, long, and reasoning-oriented video tasks, the system gains access to a gold standard for evaluating video understanding reward models.

  2. [[Long Chain-of-Thought (CoT) Reasoning Integration]]: The framework forces reward models to evaluate responses based on detailed logical steps rather than just final answers. This enables the AI system to perform complex, multi-step inference required for advanced video comprehension tasks (e.g., Explain why X happened in the video by detailing the sequence of events).

  3. [[Majority Voting Evaluation Protocol]]: Replacing single-pass or position-biased evaluation with majority voting across 8 independent judgments significantly enhances the robustness and reliability of reward signals. The improved AI system will be less susceptible to artifacts like position bias, leading to more stable and trustworthy preference learning during training.

Improved AI System Capabilities:

The resulting AI systems (VideoDRM and VideoGRM) can perform the following specific tasks:

  1. [[Robust Video Preference Judgment]]: The system can reliably distinguish between high-quality and low-quality video descriptions or generated responses by accurately predicting which response a human evaluator would prefer, even when both responses contain complex reasoning steps.

  2. [[Enhanced Reasoning Capability in Downstream Tasks]]: Because the reward models are trained on high-quality video preference data (VUP-35K), the resulting AI systems will not only judge preferences better but also exhibit improved underlying video understanding and reasoning capabilities when deployed for tasks like video question answering, temporal event detection, or complex scene analysis.

  3. [[Superior Test-Time Reranking/Selection]]: When multiple candidate answers are generated in real-time (Best-of-N settings), the system can use its reward model to select the most accurate and contextually appropriate response with high confidence, leading to superior performance in inference-time selection scenarios compared to standard Self-Judge or Majority-of-N methods.

  4. [[Cross-Modality Transfer Bridging]]: By incorporating a large corpus of image-text preference data alongside video data during training (VUP35K), the system can better generalize its reasoning skills, allowing it to transfer learned preference judgment strategies from static/textual domains to complex video understanding tasks.

Abstract

Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust evaluation benchmarks and high-quality preference data. To address this, we propose a unified framework spanning benchmark design, data construction, and reward model training. We introduce Video Understanding Reward Bench (VURB), a benchmark featuring 2,100 preference pairs with long chain-of-thought reasoning traces (averaging 1,143 tokens) and majority voting evaluation across general, long, and reasoning-oriented video tasks. We further construct Video Understanding Preference Dataset (VUP-35K) via a fully automated pipeline, providing large-scale high-quality supervision for video reward training. Building on the data, we train VideoDRM and VideoGRM, a discriminative and a generative reward model, both achieving state-of-the-art performance on VURB and VideoRewardBench. Further analysis confirms that VUP-35K enhances both reward performance and model reasoning capability, while VideoDRM and VideoGRM yield significant gains under best-of- N test-time scaling.

Sources

Related papers