Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding".
Tom: Reward models for text-to-video (T2V) generation often fail at fine-grained semantic alignment because they lack systematic verification and rely on implicit reasoning, leading to errors in handling complex prompts.
Jane: First, who's behind it and why it matters.
Paper summary: Jane: So, to wrap up our discussion on "Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding," the authors are really showing how to move beyond implicit reasoning in reward models by grounding their evaluation in an explicit plan and a structured scene graph.
Lu: It boils down to decomposing a prompt into atomic claims, verifying each one against both the video and that scene graph, and then aggregating those results using a rubric-guided analysis to get the final score. That structured approach is what they emphasize.
Meng: And they also managed to incorporate quality assessment—Visual Quality, Temporal Quality, and Physical/Common-sense Consistency—all within that same reasoning trace during the generation process. It’s a very comprehensive evaluation strategy for a single output.
Lalam: The authors acknowledge some limitations; they note that the Stage one extractor only achieves partial coverage on action-level requirements, and they also mention that the definition of Physical/Common-sense Consistency is prompt-agnostic when dealing with stylized or fantasy content.
Tom: So, while they’ve made significant progress in achieving top performance on fine-grained temporal semantics and compositional alignment, the current limitation is that the scene graph extraction doesn't perfectly cover every action detail, and they haven't fully addressed consistency across all styles yet.
Jane: It means we have a very strong method for verifying structural requirements, but we still have some areas where the AI might miss subtle details in complex creative scenarios. The paper itself is titled "Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding" and it really outlines this new framework.
Lu: The implication here is that future work will likely focus on integrating reinforcement learning to further enhance this reasoning framework, which they suggest as a path forward. That suggests the current supervised fine-tuning setup is a stepping stone.
Meng: For us in the industry, it points toward building systems where we can explicitly define what success looks like for complex video generation tasks, rather than relying on opaque reward signals. It gives us a better tool for debugging and directing model development.
Lalam: If this method continues to improve the reliability of video reward modeling, I think it could significantly elevate the quality of generated content across many domains where temporal and semantic accuracy matters immensely.
Tom: That’s what we were talking about; a more structured way to ensure that when we ask an AI for a specific type of video, it actually delivers on the detailed requirements. It’s about making the AI's internal decision-making process traceable and verifiable.
Conclusion: Tom: So, we've been diving deep into this paper, "Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding," and now it’s time for the wrap-up on what this actually means for video generation systems.
Jane: It really boils down to how they moved away from those old, fuzzy reward models by building a system that checks every single part of a prompt against visual evidence using an explicit plan and a scene graph.
Lu: That structured approach, using the spatio-temporal scene graph as that "structured visual reference," is what I find fascinating because it gives the AI something concrete to reason about beyond just raw pixels.
Meng: From my side, I’m looking at the engineering reality here—it means we can finally build reward models where we actually know *why* a video scored a certain way, not just that it did.
Lalam: I think this is huge because it allows the AI to move toward a much more reliable and consistent understanding of complex visual instructions across many different tasks.
Tom: Exactly, Jane; the title itself tells us they’ve focused on plan-and-verify reasoning powered by those scene graphs, which is a big step for tackling fine-grained semantic alignment.
Jane: And the authors really emphasize how they integrated quality checks—visual, temporal, and physical consistency—all within that single evaluation process.
Lu: The way they trained it in stages, first learning to extract those graphs and then fine-tuning it to generate the full trace, shows a really sophisticated understanding of how these components need to work together.
Meng: Practically speaking, this suggests a path forward where we can use these structured evaluations not just for training but also for debugging models that fail on specific temporal or spatial requirements.
Lalam: If we can make the AI more reliable in interpreting those detailed prompts, I see this improving how creative content is generated because the underlying structure of the video becomes much stronger.
Tom: It’s clear that by grounding their reasoning in these scene graphs and using a verification plan, they’ve created a much more transparent way for the AI to judge if a video actually meets its complex goals.
Jane: And this moves us closer to systems where we can trust the semantic alignment scores because they are backed by verifiable evidence rather than just an educated guess.
Lu: The implication is that for T2V, we’re moving toward models that don't just create aesthetically pleasing clips, but clips that accurately reflect the intricate relationships described in a prompt.
Meng: I think this opens up new avenues for testing and validating video generation pipelines where precise temporal relationships are crucial, which is a key pain point right now.
Lalam: So, moving forward, we’re looking at AI that understands the *structure* of video generation better, not just the final look.
Tom: That’s what we've been discussing; this paper sets a new standard for how reward models should approach complex semantic checks in video synthesis.
Hyomin Kim, Junghye Kim, Joanie Hayoun Chung, Yoonjin Oh, Kyungjae Lee, Sungbin Lim†, Sungwoong Kim†
Department of Artificial Intelligence, Korea University
cs.CV
Submitted: 2026-06-10
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Reward models for text-to-video (T2V) generation often fail at fine-grained semantic alignment because they lack systematic verification and rely on implicit reasoning, leading to errors in handling
Key concepts
- Spatio-Temporal Scene Graph (SG)
- This is a structured intermediate representation that organizes video information into entities, attributes, and temporal relations. It acts as a 'structured visual reference' that helps the model track event-level information across time, which is much harder for raw video data alone.
- Plan-and-Verify Reasoning
- Instead of free reasoning, the model first generates a plan to break the prompt into specific 'atomic claims' (Critical or Minor). It then verifies each claim against both the video and the scene graph, resulting in explicit outcomes like 'Supported,' 'Partially Supported,' or 'Contradicted.'
- Semantic Alignment (SA) Score
- This is a final score derived by aggregating judgments on all claims. The model writes a short analysis integrating critical and minor claim judgments to determine the overall semantic alignment, ensuring every part of the complex prompt has been addressed.
- Plan-and-Verify Structure
- This methodology involves two steps: first, creating a 'Verification Plan Generation' to decompose the prompt into specific checks. Second, performing 'SG-Grounded Claim Verification' by checking these claims against the video and scene graph to produce evidence for each judgment.
Terminology
Summary
Reward models for text-to-video (T2V) generation often fail at fine-grained semantic alignment because they lack systematic verification and rely on implicit reasoning, leading to errors in handling complex prompts. This paper proposes SG-PVR, a video reward model that addresses these limitations by employing plan-and-verify reasoning grounded in spatio-temporal scene graphs to ensure every prompt requirement is explicitly checked against visual evidence.
SG-grounded Reasoning and Scene Graph Representation
The core innovation of SG-PVR is the introduction of a spatio-temporal scene graph
(SG) as a structured intermediate representation that persists throughout the reasoning context. This graph encodes entities, attributes, and temporally grounded relations, serving as a structured visual reference.
The model extracts this graph from the video and maintains it to provide explicit evidence for verification. This approach is complementary to raw video evidence; while the video supplies pixel-level data, the scene graph organizes temporal relations and event-level information that are difficult to track with video-only reasoning.
Plan-and-Verify Reasoning Structure
SG-PVR utilizes a plan-and-verify structure
rather than freeform chain-of-thought (CoT). This involves two main steps:
-
A
Verification Plan Generation
decomposes the prompt intoatomic claims,
explicitly specifying what must be checked. These claims are tagged as eitherCritical or Minor.
-
The model performs
SG-Grounded Claim Verification,
checking each claim against both the video (V) and the scene graph (G). The model assigns one of three outcomes:Supported, Partially Supported, or Contradicted,
along with the supporting evidence drawn from G, V, or both.
Score Aggregation and Quality Evaluation
The final Semantic Alignment (SA) score is derived through a Score Aggregation
process using a predefined rubric. This rubric defines score levels based on the distribution of Supported, Partially Supported, and Contradicted outcomes.
Crucially, the model writes a short final analysis that integrates the judgments of Critical and Minor claims
to evaluate semantic alignment. In addition to SA reasoning, SG-PVR evaluates video quality along three dimensions—Visual Quality (VQ), Temporal Quality (TQ), and Physical/Common-sense Consistency (PC)—through holistic reasoning performed within the same reasoning trace,
producing both semantic alignment and quality scores in a single generation.
Training Pipeline and Performance
The model is trained in two stages. Stage 1 teaches the base VLM to extract spatio-temporal scene graphs,
using Synthetic Visual Genome 2 (SVG2) data, refining annotations to focus on semantically central
entities and relations. Stage 2 then fine-tunes the model to produce the full evaluation trace in a single generation, conditioned on a video and its prompt. The training data is constructed from diverse sources like Q-Eval-100K and VideoFeedback datasets, with score normalization applied to handle heterogeneous annotation protocols across dimensions like SA, VQ, TQ, and PC.
Experimental Validation
SG-PVR demonstrates strong performance across three evaluation settings: pointwise reward prediction on benchmarks like VSB-v2 and LGVQ; fine-grained temporal semantic understanding on a test set distinguishing videos by temporal structure (Vinoground/TimeBlind); and test-time reranking on T2V-CompBench. The results show SG-PVR consistently leads in compositional and temporal semantic alignment, achieving the highest Overall accuracy among all compared methods
for fine-grained temporal semantics, while also outperforming baselines on perceptual quality axes. Ablation studies confirm that both the scene graph grounding and the plan-and-verify structure contribute to end-to-end performance.
System Prompts for Reasoning Trace Construction
The model's reasoning is governed by specific system prompts designed to enforce structure. The Verification Plan Generation
prompt instructs the model to decompose the prompt into atomic claims, assigning a Critical or Minor
importance label and ensuring each claim begins with Verify whether.
The subsequent Semantic Reasoning Trace Generation
prompt mandates that the model evaluate each claim in order using evidence from V and G, ending with a final analysis and a score formatted as an exact output. Finally, the quality reasoning step uses system prompts to guide the model through distinct evaluations for Visual Quality (VQ), Temporal Quality (TQ), and Physical/Common-sense Consistency (PC).
Limitations
The paper acknowledges limitations, noting that the Stage 1 extractor achieves only partial coverage on action-level requirements.
Furthermore, while SG-PVR is trained with supervised fine-tuning, future work suggests integrating Reinforcement Learning to enhance its reasoning framework. The prompt-agnostic definition of Physical/Common-sense Consistency (PC) is noted as a limitation when dealing with stylized or fantasy content.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing SG-PVR, and what those improved systems will be capable of:
-
Improved Fine-Grained Semantic Alignment in Text-to-Video (T2V) Generation: The system will move beyond coarse, single-score assessments to achieve precise alignment with complex prompts.
-
Accurate Verification of Compositional Requirements: The AI will systematically decompose lengthy prompts into atomic claims and verify every single one against visual evidence, ensuring no requirement is missed or misjudged.
-
Robust Handling of Complex Temporal Semantics: The system will excel at evaluating temporal relationships, including event order, state changes (e.g., ice to water), and complex causal dependencies between actions across different time intervals within a video.
-
Enhanced Interpretability via Structured Reasoning Traces: Instead of opaque
black box
scores, the AI will generate a detailed chain-of-thought trace that explicitly shows which specific visual elements (grounded in the scene graph) supported or contradicted each prompt requirement. -
Integrated Quality Assessment: The system will produce both semantic alignment scores and intrinsic video quality scores (Visual Fidelity, Temporal Coherence, Physical/Common-sense Consistency) in a single generation, providing a holistic evaluation of the output.
-
Improved Downstream Utility for Generation: By providing highly accurate reward signals (SA + Quality), the system will serve as a superior test-time reranker or fine-tuning signal for T2V models, guiding them toward higher quality outputs based on explicit structural failures rather than uniform error averaging.
-
Differentiated Prompt Type Performance: The system will demonstrate superior accuracy across different prompt types (e.g., State Transition, Static Composition) compared to existing models, specifically by leveraging the scene graph to better handle multi-event prompts and state changes.
Abstract
Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural weaknesses in existing reasoning-based reward models: they do not systematically verify every condition described in the prompt, and the visual evidence supporting each judgment remains implicit in their free-form reasoning. We propose SG-PVR, a video reward model that addresses these limitations through plan-and-verify reasoning grounded in spatio-temporal scene graphs. The verification plan decomposes the prompt into atomic claims, making the set of requirements to be checked explicit. The spatio-temporal scene graph, encoding entities, attributes, and temporally grounded relations, is extracted from the video and maintained as a persistent structured visual reference throughout reasoning. Each claim is verified against both the video and the scene graph, anchoring judgments in explicit visual evidence. SG-PVR achieves strong performance on semantic alignment, including fine-grained temporal semantics. As a test-time reranker, it further enhances compositional alignment in T2V generation.
Sources
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
- GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
- Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
- VideoScore2: Think before You Score in Generative Video Evaluation
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
- Flow-GRPO: Training Flow Matching Models via Online RL
- Improving Video Generation with Human Feedback
- Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
- Seedance 2.0: Advancing Video Generation for World Complexity
- MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
- Wan: Open and Advanced Large-Scale Video Generative Models
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
- LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
- Unified Personalized Reward Model for Vision Generation
- Unified Reward Model for Multimodal Understanding and Generation
- DanceGRPO: Unleashing GRPO on Visual Generation
- Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models