LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
Enhuai Liu, Yunke Wang, Yutong Wang, Changming Sun, Chang Xu
The University of Sydney · CSIRO
cs.CV, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration Abstract Summary The paper introduces LoSA, a training-free sparse-attention method for accelerating video
Terminology
Summary
LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
Abstract Summary
The paper introduces LoSA, a training-free sparse-attention method for accelerating video diffusion transformers. Video diffusion transformers are costly to sample because every denoising step applies self-attention over a long 3D token sequence, with quadratic cost that dominates as resolution and duration grow. Sparse attention reduces this cost without retraining, but existing methods pursue aggressive sparsity, where further speedup costs disproportionately more attention fidelity. LoSA targets the opposite end of this trade-off: it fixes near-lossless fidelity by construction and removes as much computation as this constraint permits. Two observations make this regime practical: roughly 40% of block interactions can be removed while retaining 99% of the attention mass, and the high-mass support remains stable across denoising steps. LoSA fixes a retained-mass threshold of 99% rather than a sparsity ratio: it measures exact block attention masses at one early dense step, keeps, for each head and query block, the smallest key/value block set meeting the threshold, and reuses the frozen block indices for all remaining steps. On Wan2.1-1.3B, LoSA alone gives a 1.36× speedup with a 0.06-point VBench Overall drop. Combined with feature caching, LoSA reaches a 3.2× speedup on HunyuanVideo at a 0.02-point drop, versus 0.32 points for the strongest sparse baseline at comparable speed. Across three video diffusion transformers and speedups up to 3.2×, LoSA consistently achieves the best training-free speed–quality trade-off.
Introduction Summary
Video diffusion transformers have become a leading architecture for text-to-video generation, but their sampling cost remains high. Each denoising step applies self-attention over a long 3D token sequence, and because attention is quadratic in sequence length, its share of the sampling cost grows with resolution and duration—from roughly half of the latency at short contexts to more than 80% at long ones. Training-free acceleration is especially attractive, as it promises to lower inference cost without retraining the model, changing the sampler, or sacrificing generation quality. Existing training-free accelerators reduce cost along several complementary directions: feature caching reduces how often transformer blocks are computed by reusing or predicting intermediate outputs across denoising steps; sparse attention reduces the cost of each computed attention layer by skipping less important query–key interactions; distillation and quantization are also effective but require additional training, calibration, or model modification. This work focuses on sparse attention as a drop-in inference-time replacement for dense self-attention.
Most sparse-attention methods aim for high sparsity, pushing attention density as low as possible while keeping visual quality acceptable—SVG2, for instance, typically retains only 25–30% of blocks. This is a natural objective when sparse attention is evaluated as a standalone speedup module. However, it places the method in an aggressive approximation regime, where additional sparsity can cost a disproportionate amount of attention fidelity. The error becomes more consequential in a full acceleration pipeline, where sparse attention is combined with feature caching: cached or predicted states carry attention errors forward instead of confining them to a single computation. Thus, low attention error is not merely a conservative preference, but an important design criterion for composable inference acceleration. Simply operating an existing method at a lower sparsity level does not necessarily meet this criterion. Some methods fix a sparsity ratio, leaving the retained attention mass unspecified; others rely on coarse mass estimates that are not sufficiently accurate in the near-lossless regime.
Video diffusion attention satisfies both requirements for near-lossless sparsity. Its block-level distribution contains a long low-mass tail, allowing substantial computation to be removed with little mass loss, and after the rapidly changing early denoising steps, the high-mass support remains stable enough to be constructed once and reused throughout the remaining trajectory. These observations suggest that the right objective is not maximum sparsity, but controlled, near-lossless sparsity.
LoSA is a training-free, near-lossless replacement for self-attention in video diffusion transformers. Instead of fixing a sparsity ratio, LoSA fixes a retained-mass target. For each layer, head, and query block, it keeps the smallest set of key/value blocks whose cumulative attention mass reaches a threshold θ; the default is θ = 0.99. LoSA builds this pattern once after a few dense denoising steps and then freezes only the block indices. Because construction coincides with a dense step, the selection is driven by exact block masses rather than the coarse importance estimates that per-step sparse methods must rely on, so the retained-mass target is met accurately rather than approximately. At every subsequent step, it still recomputes Q/K/V and attention weights from the current hidden states, normalizing the softmax over the retained keys. Thus, LoSA changes only the attention support, not the model weights, denoising schedule, or cached features.
Empirically, LoSA consistently lies on the speed–quality Pareto frontier, both as a standalone module and when combined with feature caching. Its advantage over more aggressive sparse-attention configurations becomes more pronounced in cached pipelines, supporting the premise that near-lossless sparse attention is a better primitive for composable inference acceleration. This trend holds across the evaluated models and speedup levels.
The contributions are: (1) identifying and characterizing a near-lossless sparse regime in video diffusion attention, where approximately 40% of block interactions can be removed while retaining 99% of dense attention mass; (2) introducing LoSA, which constructs a retained-mass support once per sample from exact block attention masses, independently for each layer, head, and query block, and freezes only the block indices for the remaining denoising steps; (3) demonstrating that LoSA is near-lossless as a standalone module and remains effective when combined with feature caching, improving the training-free speed–quality frontier across multiple models and speedup levels.
Related Work Summary
Efficient video diffusion inference: Diffusion transformers have become a common backbone for text-to-video generation, including Wan, HunyuanVideo, and CogVideoX. Compared with image generation, video generation requires attention over much longer 3D token sequences, making inference substantially more expensive. Existing acceleration methods can be grouped by where they reduce cost: step distillation reduces the number of denoising steps; quantization reduces numerical precision and memory traffic; caching reuses intermediate features across denoising steps; sparse attention reduces the cost of attention inside a computed step. This work belongs to the last category and is training-free.
Sparse attention for video generation: Sparse attention exploits the fact that attention mass in video diffusion is highly concentrated. Some methods use predefined layouts, such as Sliding Tile Attention, Radial Attention, and DiTFastAttn, but a fixed layout may not match the content-dependent attention pattern of a given prompt or denoising step. Other methods build sparse patterns from the current attention or hidden states, such as Sparse VideoGen, Sparse VideoGen2, SpargeAttn, XAttention, Jenga, VSA, and AdaSpa. AdaSpa exploits cross-step pattern stability by reusing each searched mask for several denoising steps, but still updates the mask at preset steps and controls sparsity with a fixed budget. LoSA instead quantifies how well a single pattern preserves attention mass and shows that one pattern selected by exact retained mass can be reused for the remaining trajectory.
Feature caching: Feature caching accelerates diffusion inference by skipping part of the computation across denoising steps. Methods include DeepCache, Pyramid Attention Broadcast, Delta-DiT, FORA, ToCa, DuCa, TeaCache, FasterCache, AdaCache, and D2Cache. Caching and sparse attention are complementary: the cache decides whether a step or block is computed at all, while sparse attention lowers the cost of the attention that is still computed, so the two can be combined. However, a cache may carry attention errors from computed steps into later reused states, so the combination favors sparse attention with low error. This motivates the near-lossless design of LoSA.
Motivation Summary
The paper measures sparse attention by retained mass. For attention head h at denoising step t, let Aht = softmax(Qht Kth T / √d) be the dense attention probability matrix. Query tokens are partitioned into blocks BiQ of size R and key/value tokens into blocks BjK of size C. The block-level attention contribution from query block i to key/value block j is Mth(i, j) = Σ q∈BiQ Σ k∈BjK Aht(q, k), referred to as block attention mass. A block-sparse attention pattern keeps a subset omegahi of key/value blocks for each head h and query block i. The dense attention mass retained by this pattern is Recall t(omega) = (Σ h Σ i Σ j∈omegahi Mth(i, j)) / (Σ h Σ i Σ j Mth(i, j)). The remaining block computation is measured by Coverage(omega) = (1/H) Σ h Σ i omegahi / (nq nk), where H is the number of heads and nq, nk are the numbers of query and key/value blocks. A near-lossless sparse pattern should have recall close to one while reducing coverage meaningfully.
Near-lossless sparsity is the cost-effective regime: The relationship between retained mass and removed block area shows a strong diminishing-return effect. Removing the low-mass tail incurs little mass loss: dropping roughly 40% of the block area reduces the retained mass by only about 1%. However, pushing much further is considerably more expensive. At roughly 70% sparsity, the mass loss grows to about 7%, seven times larger. Tightening the threshold to θ = 0.999 removes only about 20% of the block area, while relaxing it to θ = 0.95 saves more computation but loses several times more mass than θ = 0.99. These results motivate optimizing retained mass directly rather than prescribing a fixed sparsity ratio. LoSA therefore fixes a retained-mass threshold, choosing θ = 0.99 inside the near-lossless region.
Near-lossless patterns are stable across denoising: At step t0 = 3, after three dense denoising steps, a block pattern that retains 99% of the block attention mass at that step is constructed and frozen. Its recall using the true dense attention mass at subsequent steps stays close to 99% throughout the remaining trajectory and is still 98.7% at the final step; even the worst-performing layers retain more than 97% of the mass. Attention starts broad and then concentrates: the top-99% region of every later step is almost a subset of the region at t0. The support contracts within an early envelope rather than migrating, so a mask built after the initial transition keeps covering the dominant regions without being updated. A natural alternative is to raise the thresholds of existing top-p estimation methods to target the same near-lossless regime, but methods such as SVG2 and SpargeAttn rely on coarse mass estimates that are not sufficiently accurate in this regime. LoSA remains close to the perfect-calibration line in the high-recall region and requires less coverage than both baselines at true recall of 99% or higher, even before accounting for their per-step estimation overhead.
Methodology Summary
LoSA replaces self-attention in video diffusion transformers without modifying model weights, prompts, sampling schedules, or classifier-free guidance settings. It operates in two stages: it constructs a mass-preserving sparse support once online for each sample and reuses the frozen support during the remaining denoising steps. No offline calibration or profiling is required. The default settings are θ = 0.99, dense attention at steps 0, 1, 2, support construction during the dense step t0 = 3, and frozen-support sparse attention afterward.
One-time pattern construction: At the construction step t0, LoSA computes the block attention masses Mth0(i, j) for every self-attention layer, head, query block, and key/value block. For each head h and query block i, let πih sort the key/value blocks by decreasing mass. Since each query row of Aht0 sums to one, the total mass of a query block of size R is R, and the retained-mass threshold θ translates into the shortest prefix shi = min m ∈ 1,..., nk: Σ r=1 m Mth0(i, πih(r)) ≥ θR. LoSA keeps omegahi = πih(1),..., πih(shi), stored independently for each layer, head, and query block, and separately for the conditional and unconditional branches under classifier-free guidance. Since omega only records block indices, its memory footprint is orders of magnitude smaller than that of a token-level mask. Because construction coincides with a dense attention step, the masses are exact rather than estimated, giving the threshold a direct meaning: setting θ = 0.99 retains 99% of the mass at t0 by construction. At later steps, near-lossless recall is maintained by the cross-step support stability.
Sparse attention with a frozen pattern: LoSA uses dense self-attention up to and including the construction step t0. Afterward, each query attends only to its retained keys. The output of head h at step t is Oth(q) = Σ k∈Kh(q) [exp(Qht(q) · Kth(k)/√d) / Σ k'∈Kh(q) exp(Qht(q) · Kth(k')/√d)] Vth(k), with Q/K/V computed from the current hidden states. Since discarded keys are removed from the softmax denominator, LoSA is not mathematically identical to dense attention; it is near-lossless in practice because omega preserves nearly all block attention mass. No auxiliary local or diagonal blocks are imposed: the retained-mass criterion directly preserves whichever blocks dominate the dense distribution. For one head, dense self-attention costs O(nq nk RCd) and frozen sparse attention costs O(Σ i omegahi RCd), where d is the head dimension. After the one-time construction step, the attention cost is reduced in proportion to the coverage.
Combination with feature caching: LoSA and feature caching act on different parts of inference. The cache decides whether a denoising step or transformer block is computed, reused, or predicted; whenever a block is computed, LoSA replaces its dense self-attention with frozen-support sparse attention. The cache threshold, reuse schedule, predictor, and cached residuals are not changed. In the combination with D2Cache, support construction is scheduled on a fully computed warm-up step, ensuring that the dense block masses required to build omega are available. After construction, skipped steps or blocks simply reuse cached states, while computed blocks use the same frozen support as in standalone inference.
Experiments Summary
Experimental setup: LoSA is evaluated on three text-to-video diffusion transformers: Wan2.1-T2V-1.3B at 480p, Wan2.1-T2V-14B at 720p, and HunyuanVideo-13B at 540p, all with their official checkpoints and default sampling configurations. Generation quality is measured with VBench on its full standard prompt suite; each configuration is evaluated with a single fixed random seed. Five representative VBench dimensions are reported (temporal flickering, dynamic degree, aesthetic quality, scene, and overall consistency), together with the aggregate Quality and Overall scores. Each method is summarized by its Overall drop ∆ relative to the corresponding dense baseline. Sparse-attention baselines are SVG1, SVG2, and SpargeAttn, all training-free. AdaSpa is not included because its code is unavailable and its reported HunyuanVideo performance is below SVG2. On the caching side, D2Cache is used, a state-of-the-art method. Methods are compared at operating points with matched end-to-end speedup: about 2.5× and 3× on Wan2.1-1.3B, 2.5× on Wan2.1-14B, and 3.2× on HunyuanVideo. Each operating point is reached by adjusting only the reuse schedule of D2Cache. All sparse-attention methods keep their default configurations throughout. All experiments run on NVIDIA H200 GPUs in a Diffusers-based inference environment. LoSA is applied to self-attention only. All layers use query block size R = 128 and key/value block size C = 32. The defaults are θ = 0.99 and t0 = 3. Block-sparse attention is implemented with FlashInfer, and the construction step uses a custom dense-attention kernel that accumulates exact block masses while computing full attention. Construction adds a one-time overhead of roughly one denoising step, which is included in all reported latencies and amortized over the remaining sparse steps.
Main results: On Wan2.1-1.3B, LoSA alone accelerates sampling by 1.36× while reducing VBench Overall by only 0.06 points, staying close to dense attention on every reported dimension. SVG2 reaches a higher standalone speedup (1.69×) but loses 0.45 points, over seven times the degradation. SVG1 (1.75×) and SpargeAttn (1.30×) lose 2.01 and 2.84 points, and SpargeAttn is even slower than LoSA. LoSA deliberately trades part of the standalone speedup for near-losslessness.
Under composition with feature caching, at the ≈ 2.5× operating point, D2Cache alone reaches 2.33× with a 0.15-point drop. Composing it with SVG2 raises the speedup to 2.50× but increases the quality loss to 0.61 points, worse than the cache alone: the attention error of the sparse module is written into cached states and propagated across reused steps. Composing the same cache with LoSA reaches the same 2.50× with the quality drop unchanged at 0.15 points. At ≈ 3×, LoSA+D2Cache is simultaneously the fastest (3.09×) and the highest-quality (∆ = −0.30) configuration, surpassing both D2Cache alone (2.84×, −0.41) and SVG2+D2Cache (2.92×, −0.71).
On Wan2.1-14B at 720p, where a dense sample takes roughly 2000 seconds, LoSA+D2Cache achieves 2.50× with a 0.43-point drop, against 2.46× and 0.71 points for SVG2+D2Cache. On HunyuanVideo-13B, at a matched ≈ 3.2× speedup, LoSA+D2Cache is virtually lossless (∆ = −0.02) while SVG2+D2Cache loses 0.32 points. The near-lossless regime is thus not specific to a single model family.
Ablation studies: The ablations examine the two design choices of LoSA: the selection rule and the retained-mass threshold. Both act directly on retained attention mass, so they are evaluated by recall rather than benchmark scores. Masses are recorded on Wan2.1-1.3B over 100 prompts across all layers, heads, and denoising steps; latencies are measured end-to-end on the same prompts.
Retained mass vs. a fixed sparsity ratio: Some block-sparse methods use a fixed sparsity level shared across heads. However, attention concentration varies widely from head to head: some heads concentrate their mass on a few blocks, while others spread it broadly. LoSA selects blocks by cumulative mass instead of a fixed ratio. A fixed-ratio top-k variant matched to the same average coverage runs equally fast but retains only 96.5% of the mass, more than three times the default’s loss, and its worst-layer recall falls to 86.6% against 97.3%: the uniform budget truncates exactly those heads whose attention is spread broadly.
Choice of θ: The realized recall never falls more than 0.1 points below the target (and sits above it at θ = 0.95), so θ is a reliable fidelity control. The default follows from the sharply nonlinear trade between recall and time: giving up the first point of mass (θ = 0.99) saves 28.5 seconds, while giving up three more points (θ = 0.95) recovers only 7.3 more. The default is therefore θ = 0.99, the knee of this trade-off.
Conclusion Summary
The paper identified a near-lossless sparse regime in video diffusion attention: roughly 40% of block interactions can be removed while retaining 99% of the attention mass, and the high-mass support is stable enough to be constructed once and reused. LoSA exploits this regime with a deliberately simple design: one dense step yields exact block masses, the threshold θ specifies the retained mass directly, and every subsequent step reuses the frozen block indices. As a standalone module, LoSA delivers a meaningful speedup while remaining nearly lossless. Composed with feature caching, it multiplies the speedup of a strong cache at almost no additional quality cost, whereas the aggressive sparse-attention baseline degrades the combination below the cache alone. The result is the best training-free speed–quality trade-off across all evaluated models and operating points. The authors hope near-lossless sparse attention becomes a standard primitive for composable diffusion inference acceleration.
Improvements for AI systems
Based on this paper, here are specific improvements I can make to AI systems and what the improved systems can do:
-
Improvement: Implement LoSA as a drop-in replacement for dense self-attention in video diffusion transformers, using a fixed 99% retained-mass threshold rather than a fixed sparsity ratio.
-
Capability: The improved system can accelerate video generation by 1.36× with only a 0.06-point quality drop (VBench Overall), versus 0.45-point drops for aggressive sparse methods at similar speeds.
-
Improvement: Build the sparse attention pattern once at denoising step 3 (after 3 dense steps) and freeze the block indices for all remaining steps, leveraging the observation that high-mass attention support remains stable (98.7% recall at final step).
-
Capability: The system eliminates per-step mask estimation overhead while maintaining near-lossless fidelity, reducing computational cost without repeated pattern searches.
-
Improvement: Replace coarse importance estimates (used by SVG2, SpargeAttn) with exact block attention masses computed during a dense step, ensuring the retained-mass threshold is met precisely rather than approximately.
-
Capability: The system achieves perfect calibration in the high-recall region, requiring less coverage than baselines at 99% recall while guaranteeing the specified mass retention.
-
Improvement: Integrate LoSA with D2Cache (feature caching) without modifying cache thresholds, reuse schedules, or predictors—only replacing dense attention in computed blocks with frozen sparse attention.
-
Capability: The combined system reaches 3.2× speedup on HunyuanVideo with only a 0.02-point quality drop (versus 0.32 points for aggressive sparse baselines), and 3.09× on Wan2.1-1.3B with a 0.30-point drop—outperforming both cache-alone and cache+aggressive-sparse configurations.
-
Improvement: Use cumulative mass-based selection independently for each attention head and query block, rather than a uniform sparsity ratio across heads.
-
Capability: The system preserves 99% mass even for heads with broadly spread attention (worst-layer recall 97.3% vs 86.6% for fixed-ratio top-k), avoiding truncation of diffuse attention patterns.
-
Improvement: Expose the retained-mass threshold θ (default 0.99) as the primary control parameter, with realized recall never falling more than 0.1 points below target.
-
Capability: Users can precisely trade off speed and quality—e.g., θ=0.95 saves 35.8 seconds total but loses 4× more mass than θ=0.99, enabling informed configuration for different latency/quality requirements.
-
Improvement: Apply LoSA without retraining, weight modification, sampler changes, or offline calibration—works across Wan2.1-1.3B, Wan2.1-14B, and HunyuanVideo-13B.
-
Capability: The system provides immediate inference acceleration for any existing video diffusion transformer, with consistent Pareto-optimal speed-quality trade-offs across model families and resolutions (480p to 720p).
-
Improvement: Store only block indices (not token-level masks) for the frozen sparse pattern, with memory footprint orders of magnitude smaller than token-level alternatives.
-
Capability: The system maintains low memory overhead even for long 3D token sequences, enabling deployment on memory-constrained hardware without sacrificing acceleration benefits.
Sources
- $\Delta$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
- QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
- FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
- FORA: Fast-Forward Caching in Diffusion Transformer Acceleration
- Wan: Open and Advanced Large-Scale Video Generative Models
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
- Improved Distribution Matching Distillation for Fast Image Synthesis
- DiTFastAttn: Attention Compression for Diffusion Transformer Models
- VSA: Faster Video Diffusion with Trainable Sparse Attention
- Training-Free Efficient Video Generation via Dynamic Token Carving
- Real-Time Video Generation with Pyramid Attention Broadcast
- Accelerating Diffusion Transformers with Token-wise Feature Caching
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models