PhysPlan: Grounded Physical State Reasoning and Graph-Guided Optimization for Physically Plausible Video Generation

arXiv:2609.13006 · cs.CV · Submitted 2026-07-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "PhysPlan: Grounded Physical State Reasoning and Graph-Guided Optimization for Physically Plausible Video Generation".

Jane: Video diffusion models (VDMs) have shown remarkable success in generating high-fidelity video content, but they fundamentally lack an intrinsic understanding of physical laws,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re starting with "PhysPlan: Grounded Physical State Reasoning and Graph-Guided Optimization for Physically Plausible Video Generation." That title tells us immediately that they are combining logical reasoning with optimization techniques to get physical consistency in video generation.

Jane: It really does sound like a plan; they aren't just throwing physics in there randomly, but they’re building a system where the visual understanding guides the physical simulation.

Lu: The authors are from several institutions, which suggests a strong interdisciplinary effort to bridge computer vision and physical modeling concepts together.

Meng: I see that they mention coupling Vision-Language Models with test-time optimization; that means the VLM does the thinking part, and then the optimization step makes sure the visuals follow those thoughts physically.

Lalam: That structure suggests a significant leap because it moves away from just hoping the visual output looks right, toward actively enforcing causal dynamics during generation.

Tom: It’s about taking that high-level planning done by an AI and forcing it to obey real-world rules like gravity or fluid dynamics when it renders the video.

Jane: That sounds like a much more robust approach than just relying on the model's general knowledge of what things look like.

Lu: The concept of agentic physical planning, where the VLM constructs a scene graph and predicts discrete physical events called the Physical Delta, is quite sophisticated.

The paper's summary: Tom: So, to summarize what PhysPlan actually does for us, it proposes a two-stage framework that uses a Vision-Language Model to plan the physics first, and then uses that plan to guide the diffusion process during generation.

Jane: It’s essentially using the VLM as a cognitive simulator that figures out the sequence of physical changes before anything is even drawn, which is pretty smart because it sets up the entire timeline correctly.

Meng: The paper highlights two key technical additions: Object-Centric Gradient Routing to stop background drift, and Kinetic Intensity Profiling to dynamically adjust how aggressively we optimize during high-change moments.

Lalam: Those dynamic adjustments are crucial because they allow the system to be flexible enough to handle things like a sudden shatter or a slow melt without needing a fixed schedule that might fail.

Lu: They ground the visual states using methods like Kinematic Masking and Geometric Depth Extraction, which gives the optimization process concrete spatial and physical constraints to work against.

Tom: The paper shows that when they test it on benchmarks like PhyGenBench and Physics-IQ, PhysPlan does a much better job than existing models at keeping things physically consistent across time.

Jane: So, the main takeaway is that this method significantly improves the physical plausibility and temporal coherence of generated videos by enforcing real-world causal dynamics through this planning and optimization loop.

Meng: That focus on maintaining structural integrity during complex motion seems like a very practical win for anyone needing realistic simulations in AI applications.

The paper's improvements: Tom: Now let’s talk about the specific improvements they introduced, because that’s where the real technical meat of PhysPlan lies. They suggest coupling VLM-driven planning with training-free optimization for a better outcome.

Jane: They focus on several distinct enhancements, including Object-Centric Gradient Routing to keep the background stable and Kinetic Intensity Profiling to modulate learning rates based on physical severity.

Lu: The kinetic intensity profiling is particularly interesting because it allows the guidance dynamics to adapt; instead of using a uniform schedule, it changes based on how volatile the physical state is, which static schedules often miss.

Meng: From an implementation standpoint, that dynamic modulation sounds powerful but also complex to tune correctly so it doesn't cause instability during the optimization phase.

Lalam: The idea of using the VLM to create a "Chain-of-Visual-Thought" representation of the event before generating frames is a major conceptual improvement for how we think about AI agents creating content.

Tom: They also detail how they use Efficient Manifold Projection via Latent Slicing during Stage two optimization to handle computational constraints when processing those K target frame indices established in Stage one.

Jane: So, they’ve managed to create a system that handles complex physical states by breaking the problem into discrete, manageable planning steps guided by physical parameters.

Conclusion: Tom: So, wrapping up this discussion on PhysPlan: the authors have shown how coupling agentic physical planning with training-free optimization can enforce causal dynamics in video generation.

Jane: It seems the core finding is that injecting physical awareness via a two-stage process—planning and then guided optimization—significantly boosts the real-world adherence of generated videos compared to previous methods.

Lu: The implication for us is that we might be able to move toward generative AI that doesn't just mimic appearance but actually simulates realistic physical behavior across sequences.

Meng: For practical use, this suggests a path toward creating more reliable simulation tools where the output isn't just visually convincing but also structurally sound according to physical laws.

Lalam: I think the most profound impact is on culture because if we can reliably generate physically plausible complex scenes, it opens up entirely new avenues for creative and scientific visualization that are currently inaccessible to diffusion models.

Tom: Absolutely, PhysPlan offers a solid blueprint for moving past mere visual aesthetics toward systems that respect the underlying physics of reality when they create video.

Minh-Loi Nguyen, Xuan-Vu Le, Thanh-Toan Do, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

University of Science, Ho Chi Minh City, Vietnam · Vietnam National University, Ho Chi Minh City, Vietnam · Monash University, Melbourne, Victoria, Australia · University of Dayton, Dayton, Ohio

cs.CV

Submitted: 2026-07-16

Updated: 2026-09-29

Project page: https://physplan.github.io

Importance score: 92/100

The gist: Video diffusion models (VDMs) have shown remarkable success in generating high-fidelity video content, but they fundamentally lack an intrinsic understanding of physical laws, often resulting in

Key concepts

PhysPlan
A framework that combines Vision-Language Models with graph-guided optimization to achieve physically plausible video generation by enforcing causal dynamics through planning and optimization steps.
Agentic Physical Planning
The process where a Vision-Language Model constructs a scene graph and predicts discrete physical events called the Physical Delta, allowing the AI to plan the sequence of physical changes before drawing frames.
Object-Centric Gradient Routing
A technical addition that helps stop background drift during optimization by routing gradients specifically based on objects, which keeps the background stable during generation.
Kinetic Intensity Profiling
A technique that dynamically adjusts how aggressively the system optimizes its guidance dynamics. It changes the optimization schedule based on how volatile or severe a physical state change is.

Terminology

Summary

Video diffusion models (VDMs) have shown remarkable success in generating high-fidelity video content, but they fundamentally lack an intrinsic understanding of physical laws, often resulting in causally illogical sequences and structural hallucinations. This paper introduces PhysPlan, a novel training-free guidance framework that shifts the paradigm from stochastic visual interpolation to agentic physics simulation by coupling the logical reasoning of Vision-Language Models (VLMs) with test-time optimization. By injecting physical awareness via a two-stage process involving cognitive planning and object-centric gradient routing, PhysPlan aims to enforce real-world causal dynamics, significantly improving the physical plausibility and temporal coherence of generated videos on benchmarks like PhyGenBench and Physics-IQ.

How it works

PhysPlan operates in two distinct stages: Agentic Physical Planning (Stage 1) and Training-Free Physics Guidance (Stage 2). Stage 1 utilizes a VLM, such as Gemini, to act as an iterative cognitive simulator. This stage decomposes multimodal inputs into a Chain-of-Visual-Thought to construct a structured representation of the physical event. The VLM constructs a Scene Graph G0 = (V0, E0) representing entities and their relationships, and infers the governing physical laws to predict a sequence of discrete transitions called the Physical Delta (∆phys) = Σδi. This planning is not uniform; instead, it iterates strictly over these K events defined in ∆phys. For each event, the VLM assigns a temporal anchor using a progress ratio ri ∈ (0, 1], mapping this to a target frame index fi = ⌊ri × F⌋. The intermediate visual keyframe Ii is then synthesized by applying a conditional generative prior, specifically the Nano Banana 2 instruction-guided image editing model Ψ, to render severe non-linear physical changes.

Multimodal Grounding and Kinetic Intensity Profiling

To translate the generated visual states into mathematically enforceable conditions for optimization, PhysPlan employs multimodal grounding. This involves two key mechanisms:

  1. Kinematic Masking (Mi): Grounded-SAM-2 is applied to Ii using active node labels from Gi to extract precise binary masks, isolating active subjects from the passive environment.

  2. Geometric Depth Extraction (Di): DepthAnythingV2 is used to extract a high-resolution monocular depth map, establishing the 3D spatial layout and physical volume boundaries.

Crucially, PhysPlan introduces Kinetic Intensity Profiling, which dynamically parameterizes framework hyperparameters based on physical severity. Instead of uniform updates, the VLM evaluates the ordered sequence of physical states G = Σ(G0, ri) to extract a normalized Kinetic Intensity Profiling w ∈ [0, 1]F. This profile is then used to modulate guidance dynamics:

- Adaptive Learning Rate Modulation (ηf): The base step size ηbase is scaled as ηf = ηbase · (1 + λwf), which forcefully increases the gradient updates specifically during high-entropy events.

- Guidance Step Density: During "highly volatile windows (where wf approaches 1.0), we dynamically increase the number of localized gradient iterations M, ensuring the latent trajectory strictly adheres to the complex topological changes predicted by the agentic planning phase."

Training-Free Graph-Guided Video Optimization

Stage 2 involves using the VDM as a fine-grained motion synthesizer guided by Stage 1's outputs. To handle computational constraints, PhysPlan employs Efficient Manifold Projection via Latent Slicing. For the K target frame indices established in Stage 1, it extracts only their corresponding temporally local latent slices J = ΣjK and estimates the fully denoised target frames efficiently using Tweedie’s formula: Δx0t fi = D(z0t ji).

The optimization objective is a unified loss function Ltotal = Lspatial + Ldepth, computed sparsely at the K temporally anchored keyframes.

  1. Masked Spatial Trajectory Loss (Lspatial): This term ensures adherence to the synthesized kinematic trajectories:

- The loss is computed purely on active subjects: Lspatial = (1/K) Σ [Mfi ⊙ (Δx0t fi - I fi) 22].

  1. 3D Structural Consistency Loss (Ldepth): This constrains the depth geometry against extracted priors:

- The loss is computed against the DeepAnythingV2 model: Ldepth = (1/K) Σ Ddepth(Δx0t fi) - Df i 22.

Object-Centric Gradient Routing and Adaptive VLO

To prevent background drift, PhysPlan introduces Object-Centric Gradient Routing. This mechanism explicitly masks the backpropagated gradients, enforcing an absolute environment lock by routing the gradient exclusively through the active region defined by downsampled masks Mlatent fi.

Improvements for AI systems

Here are specific improvements that can be made to existing Video Diffusion Models (VDMs) by incorporating the PhysPlan framework, and what these improved AI systems will be capable of:


The following improvements stem from integrating the PhysPlan framework into VDM architectures, moving them from stochastic interpolators to physically aware simulators.

  1. Improved Physical Plausibility and Causal Reasoning:

A VDM can generate videos where object motions rigorously adhere to real-world physics (e.g., gravity, fluid dynamics, rigid body mechanics) and maintain perfect object permanence across long sequences.

  1. Elimination of Structural Hallucinations: The system will no longer produce hallucinated or illogical transitions (e.g., liquids defying viscosity, objects morphing illogically). Instead, it will enforce physically plausible state changes dictated by the input prompt and causal reasoning.

  2. Preservation of Passive Backgrounds: The optimized video generation process will completely lock the passive environment during complex motion, preventing background drift, temporal flickering, or layout degradation that plague current global optimization methods.

  3. High-Fidelity Topological Deformation Handling: The system can accurately synthesize complex structural changes (e.g., shattering glass, melting ice) by dynamically adjusting guidance intensity based on the severity of the physical deformation (Kinetic Intensity Profiling), rather than relying on rigid, uniform schedules that fail in heterogeneous events.

  4. Agentic Physics-Driven Planning: The AI system will utilize a Vision-Language Model (VLM) as an iterative cognitive simulator to decompose complex prompts into a structured Chain-of-Visual-Thought and predict the necessary sequence of physical events (Physical Delta). This allows the system to plan non-linear, multi-step physical sequences before generation begins.

  5. Dynamic Guidance Parameter Adaptation: The model will adapt its internal diffusion sampling hyperparameters (learning rates and guidance step density) in real-time during inference based on a computed Kinetic Intensity Profile, ensuring that high-entropy physical transitions receive the necessary intensive structural enforcement to maintain coherence.

In summary, the improved AI system will transition from merely synthesizing aesthetically pleasing visual patterns to acting as an autonomous agent capable of generating videos that are not just photorealistic, but are fundamentally grounded in and respect the laws of physics.

Sources

Related papers