SubjectAnchor: Subject-Aware Memory-to-Video for Multi-Shot Storytelling
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- Cosmos World Foundation Model Platform for Physical AI
- Wan: Open and Advanced Large-Scale Video Generative Models
- PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
- StreamingEffect: Real-Time Human-Centric Video Effect Generation
- VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Controllable Video Generation: A Survey
- A Survey on Long Video Generation: Challenges, Methods, and Prospects
- VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
- SkyReels-A2: Compose Anything in Video Diffusion Transformers
- From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
- 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
- HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- ShoulderShot: Generating Over-the-Shoulder Dialogue Videos
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- Helios: Real Real-Time Long Video Generation Model
- StoryMem: Multi-shot Long Video Storytelling with Memory
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
- MoCha: Towards Movie-Grade Talking Character Synthesis
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models