FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Project page: https://yinbo0927.github.io/FrameMorrow
Terminology
Sources
- Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
- LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
- LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation
- MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
- FrameOracle: Learning What to See and How Much to See in Videos
- Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
- ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
- SkillGraph: Self-Evolving Multi-Agent Collaboration with Multimodal Graph Topology
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- Memento: Reconstruct to Remember for Consistent Long Video Generation
- LongLive: Real-time Interactive Long Video Generation
- DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
- Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
- FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
- SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
- Dual Latent Memory for Visual Multi-agent System
- StoryMem: Multi-shot Long Video Storytelling with Memory
- EgoLCD: Egocentric Video Generation with Long Context Diffusion
- QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models