PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
cs.CV, cs.AI, cs.GR
Submitted: 2026-09-15
Updated: 2026-09-21
Project page: https://czzzzh.github.io/PhysStream
License: http://creativecommons.org/licenses/by/4.0/
The gist: Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes.
Terminology
Abstract
Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the full control schedule before generation starts, or use pixel-space signals that dictate object positions rather than physical dynamics. To address these limitations, we propose PhysStream, an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory---positional maps and object tracking maps derived online from previously generated frames---and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics. We train our model in two stages: a bidirectional model is first finetuned with motion-control conditioning, then a causal autoregressive model is trained with additional structured scene memory, further improving physical consistency. PhysStream enables interactive, mid-generation control over multi-object tabletop rigid-body scenes---a capability not supported by prior methods---reducing motion distribution distance (FVMD) by 33% and trajectory error by 12% over the strongest baselines on synthetic benchmarks, and is preferred by human evaluators in over 85% of in-the-wild comparisons. Please check our website for more details: https://czzzzh.github.io/PhysStream
Sources
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Physical Simulator In-the-Loop Video Generation
- 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
- Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
- Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Pyramidal Flow Matching for Efficient Video Generative Modeling
- Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
- FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
- Depth Anything 3: Recovering the Visual Space from Any Views
- Flow Matching for Generative Modeling
- Fr'echet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- RealWonder: Real-Time Physical Action-Conditioned Video Generation
- Do generative video models understand physical principles?
- SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation
- SAM 2: Segment Anything in Images and Videos
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models