Beyond Inpainting: Unleash 3D Understanding for Stable Camera-Controlled Video Generation
cs.CV, cs.GR
Submitted: 2026-01-15
Updated: 2026-09-27
Project page: https://eleanor6725.github.io/DepthDirector
Terminology
Sources
- AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
- SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
- GS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
- Wan: Open and Advanced Large-Scale Video Generative Models
- Dens3R: A Foundation Model for 3D Geometry Prediction
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- Training-free Camera Control for Video Generation
- LoRA: Low-Rank Adaptation of Large Language Models
- ViPE: Video Pose Engine for 3D Geometric Perception
- FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction
- Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
- MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
- Depth Anything 3: Recovering the Visual Space from Any Views
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models