Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
cs.LG
Submitted: 2026-09-17
Updated: 2026-09-21
Comments: 20 pages, 10 figures
Code: https://github.com/OpenVDN/vdn-minimax-h3
Project page: https://openvdn.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck.
Terminology
Abstract
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which updates memory once per frame by jointly incorporating its spatial tokens. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5x speedup over the 50-step dense H3 baseline on the same GPU count.
Sources
- Qwen3-VL Technical Report
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
- PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference
- xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
- Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- LTX-Video: Realtime Video Latent Diffusion
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- Pyramidal Flow Matching for Efficient Video Generative Modeling
- Kimi Linear: An Expressive, Efficient Attention Architecture
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- LLM In-Context Recall is Prompt Dependent
- VideoMamba: State Space Model for Efficient Video Understanding
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- Movie Gen: A Cast of Media Foundation Models
- Adversarial Diffusion Distillation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks