Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
cs.CV, cs.AI, cs.RO
Submitted: 2026-03-03
Updated: 2026-08-30
Comments: v2:Expanded the experiment section with more baselines and add more experiments in supplementary--corrected some typographical errors, and corrected author-affiliation information that was inaccurate in the previous version
Project page: https://sensational-brioche-7657e7.netlify.app
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models.
Terminology
Abstract
Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time. In this work, we present Phys4D, a pipeline for learning physics-consistent 4D world representations from video diffusion models. Phys4D adopts a three-stage training paradigm that progressively lifts appearance-driven video diffusion models into physics-consistent 4D world representations. We first bootstrap robust geometry and motion representations through large-scale pseudo-supervised pretraining, establishing a foundation for 4D scene modeling. We then perform physics-grounded supervised fine-tuning using simulation-generated data, enforcing temporally consistent 4D dynamics. Finally, we apply simulation-grounded reinforcement learning to correct residual physical violations that are difficult to capture through explicit supervision. To evaluate fine-grained physical consistency beyond appearance-based metrics, we introduce a set of 4D world consistency evaluation that probe geometric coherence, motion stability, and long-horizon physical plausibility. Experimental results demonstrate that Phys4D substantially improves fine-grained spatiotemporal and physical consistency compared to appearance-driven baselines, while maintaining strong generative performance. Our project page is available at https://sensational-brioche-7657e7.netlify.app/
Sources
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
- Training Diffusion Models with Reinforcement Learning
- TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model
- TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement
- LoRA: Low-Rank Adaptation of Large Language Models
- Inference of Time-Reversal Asymmetry from Time Series in a Piezoelectric Energy Harvester
- VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
- CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
- Patchview: LLM-Powered Worldbuilding with Generative Dust and Magnet Visualization
- SoftMAC: Differentiable Soft Body Simulation with Forecast-based Contact Model and Two-way Coupling with Articulated Rigid Bodies and Clothes
- RealWonder: Real-Time Physical Action-Conditioned Video Generation
- FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation
- Calculating HOMFLY-PT polynomials on a photonic processor
- Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis
- Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
- Do generative video models understand physical principles?
- T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
- Movie Gen: A Cast of Media Foundation Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models