Bootstrapping Video Interaction Generation with Synthetic State Transitions
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/black-forest-labs/flux
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Intuitive physics understanding emerges from self-supervised pretraining on natural videos
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- Classifier-Free Diffusion Guidance
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
- Diffusion Models for Video Prediction and Infilling
- GPT-4o System Card
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models
- Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
- Flow Matching for Generative Modeling
- Large Language Models have Intrinsic Self-Correction Ability
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency
- Latte: Latent Diffusion Transformer for Video Generation
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models