Improving Video Sparse Attention with Fine-grained Router and Sparse Rebasing
cs.CV
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Mixture of Contexts for Long Video Generation
- Training Deep Nets with Sublinear Memory Cost
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
- SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
- Seedance 1.0: Exploring the Boundaries of Video Generation Models
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
- Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
- Open-Sora Plan: Open-Source Large Video Generation Model
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
- Movie Gen: A Cast of Media Foundation Models
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- Progressive Distillation for Fast Sampling of Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models