Fluid-Gen-Zero: Grounding Pretrained Video Generators in Physics without Training
cs.CV, cs.GR
Submitted: 2026-10-07
Updated: 2026-10-07
Code: https://github.com/alimamacreative/FLUX-Controlnet-Inpainting
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
- Physical Simulator In-the-Loop Video Generation
- PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
- MotionClone: Training-Free Motion Cloning for Controllable Video Generation
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- RealWonder: Real-Time Physical Action-Conditioned Video Generation
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
- SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation
- SAM 2: Segment Anything in Images and Videos
- Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising
- ATI: Any Trajectory Instruction for Controllable Video Generation
- WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
- PhysAnimator: Physics-Guided Generative Cartoon Animation
- PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models