Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
- LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
- Seedance 1.0: Exploring the Boundaries of Video Generation Models
- EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
- LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation
- Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
- ViMax: Agentic Video Generation
- Kling-Omni Technical Report
- GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
- ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
- ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
- Memento: Reconstruct to Remember for Consistent Long Video Generation
- Automated Movie Generation via Multi-Agent CoT Planning
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
- EvoState: Closed-Loop Visual State Management for Long-Form Video Generation
- StoryMem: Multi-shot Long Video Storytelling with Memory
- VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
- VideoMemory: Toward Consistent Video Generation via Memory Integration
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models