Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
cs.CV, cs.AI, cs.LG
Submitted: 2026-01-22
Updated: 2026-08-26
Code: https://github.com/hpcaitech/Open-Sora
Project page: https://dohunlee1.github.io/MemoryV2V/1
Terminology
Sources
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
- Mixture of Contexts for Long Video Generation
- LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
- Goku: Flow Based Video Generative Foundation Models
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- CCEdit: Creative and Controllable Video Editing via Diffusion Models
- TokenFlow: Consistent Diffusion Features for Consistent Video Editing
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
- Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
- LoViC: Efficient Long Video Generation with Context Compression
- VACE: All-in-One Video Creation and Editing
- LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- Auto-Encoding Variational Bayes
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models