The Past Frames the Future: Memory for Autoregressive Video Generation
cs.CV
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/HaroldChen19/Awesome-AR-Video-Memory
Terminology
Sources
- From Zero to Hero: Training-Free Custom Concept Spawning in World Models
- Direct Motion Models for Assessing Generated Videos
- Differentiable Neural Computers with Memory Demon
- DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis
- Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
- CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
- Titans: Learning to Memorize at Test Time
- Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
- Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
- What-If World: A Causal Benchmark for General World Models in Embodied Scenarios
- Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation
- 3DSPA: A 3D Semantic Point Autoencoder for Evaluating Video Realism
- Towards Error-Free Long Video Generation
- PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
- SkyReels-V2: Infinite-length Film Generative Model
- Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation
- Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models