Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher
cs.CV, cs.AI, cs.LG
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- Video Diffusion Models
- Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- LongLive: Real-time Interactive Long Video Generation
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- SkyReels-V2: Infinite-length Film Generative Model
- MAGI-1: Autoregressive Video Generation at Scale
- World Models
- DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
- Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
- Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation
- Score-Based Generative Modeling through Stochastic Differential Equations
- Scalable Diffusion Models with Transformers
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models