4Director: Controlling Video World Models with Rigid 3D Geometry
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://stability-ai.github.io/4director
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
- SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
- Qwen3-VL Technical Report
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Pseudo-Simulation for Autonomous Driving
- SAM 3: Segment Anything with Concepts
- DeepVerse: 4D Autoregressive Video Generation as a World Model
- Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
- LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
- I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
- I2VControl: Disentangled and Unified Video Motion Synthesis Control
- World Models
- Mastering Diverse Domains through World Models
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- RELIC: Interactive Video World Model with Long-Horizon Memory
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models