Phase-aware video generation for physics-grounded dynamics and interactions
cs.CV
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- PhysGen3D: Crafting a Miniature Interactive World from a Single Image
- PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
- Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
- NEWTON: Agentic Planning for Physically Grounded Video Generation
- 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
- PhysPlan: Grounded Physical State Reasoning and Graph-Guided Optimization for Physically Plausible Video Generation
- Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision
- PhysMotion: Physics-Grounded Dynamics From a Single Image
- Neural Latent Arbitrary Lagrangian-Eulerian Grids for Fluid-Solid Interaction
- Wan: Open and Advanced Large-Scale Video Generative Models
- Learning Lagrangian Fluid Mechanics with E($3$)-Equivariant Graph Neural Networks
- SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion
- ATI: Any Trajectory Instruction for Controllable Video Generation
- Streaming Video Generation with Streaming Force Control
- ObjCtrl-2.5D: Training-free Object Control with Camera Poses
- HunyuanVideo 1.5 Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models