WorldAttention: An Efficient Attention Architecture for Interactive Video World Models
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/alibaba-damo-academy/WorldAttention
Project page: https://alibaba-damo-academy.github.io/WorldAttention
Terminology
Sources
- Qwen Technical Report
- Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
- Scaling Linear Attention with Sparse State Expansion
- Mixture of Contexts for Long Video Generation
- SkyReels-V2: Infinite-length Film Generative Model
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- Autoregressive Video Generation without Vector Quantization
- StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
- Long Context Tuning for Video Generation
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- SEA: Sparse Linear Attention with Estimated Attention Mask
- FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
- MAGI-1: Autoregressive Video Generation at Scale
- Wan: Open and Advanced Large-Scale Video Generative Models
- Linformer: Self-Attention with Linear Complexity
- Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
- RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models