Lifelong Learning of Video Diffusion Models From a Single Video Stream
cs.CV, cs.LG
Submitted: 2024-06-07
Updated: 2026-09-22
Comments: Video samples are available here: https://drive.google.com/drive/folders/1CsmWqug-CS7I6NwGDvHsEN9FqN2QzspN
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- Diffusion for World Modeling: Visual Details Matter in Atari
- Lumiere: A Space-Time Diffusion Model for Video Generation
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- On Tiny Episodic Memories in Continual Learning
- Semantically Consistent Video Inpainting with Conditional Diffusion Models
- Learning from Streaming Video with Orthogonal Gradients
- Diffusion Reward: Learning Rewards via Conditional Video Diffusion
- CCD: Continual Consistency Diffusion for Lifelong Generative Modeling
- Locality Sensitive Sparse Encoding for Learning World Models Online
- Orca: Progressive Learning from Complex Explanation Traces of GPT-4
- Progressive Neural Networks
- ViLCo-Bench: VIdeo Language COntinual learning Benchmark
- LLaMA: Open and Efficient Foundation Language Models
- Online Curvature-Aware Replay: Leveraging $\mathbf{2^{nd}}$ Order Information for Online Continual Learning
- Continual Learning and Catastrophic Forgetting
- Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment
- Exploring Continual Learning of Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models