Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation
cs.CV
Submitted: 2026-09-17
Updated: 2026-09-17
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- MAGI-1: Autoregressive Video Generation at Scale
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- SkyReels-V2: Infinite-length Film Generative Model
- Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- LoL: Longer than Longer, Scaling Video Generation to Hour
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Autoregressive Video Generation without Vector Quantization
- Flex Attention: A Programming Model for Generating Optimized Attention Kernels
- End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
- LTX-Video: Realtime Video Latent Diffusion
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- FIFO-Diffusion: Generating Infinite Videos from Text without Training
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
- Train Short, Inference Long: Training-free Horizon Extension for Autoregressive Video Generation
- Unified Video Action Model
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models