Parameterized Stripe Attention for Efficient Video Generation
cs.CV, cs.AI
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/black-forest-labs/flux
Terminology
Sources
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Flow Matching for Generative Modeling
- Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers
- RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Denoising Diffusion Implicit Models
- ThunderKittens: Simple, Fast, and Adorable AI Kernels
- VORTA: Efficient Video Diffusion via Routing Sparse Attention
- Wan: Open and Advanced Large-Scale Video Generative Models
- ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
- VMoBA: Mixture-of-Block Attention for Video Diffusion Models
- Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
- Training-free and Adaptive Sparse Attention for Efficient Long Video Generation
- Why Attention Patterns Exist: A Unifying Temporal Perspective Analysis
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
- FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
- Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models