SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
cs.CV, cs.LG
Submitted: 2026-02-27
Updated: 2026-09-20
Code: https://github.com/vita-epfl/SenCache
License: http://creativecommons.org/licenses/by/4.0/
The gist: Diffusion models achieve state-of-the-art video generation quality, but their inference remains expensive due to the large number of sequential denoising steps.
Terminology
Abstract
Diffusion models achieve state-of-the-art video generation quality, but their inference remains expensive due to the large number of sequential denoising steps. This has motivated a growing line of research on accelerating diffusion inference. Among training-free acceleration methods, caching reduces computation by reusing previously computed model outputs across timesteps. Existing caching methods rely on heuristic criteria to choose cache/reuse timesteps and require extensive tuning. We address this limitation with a principled sensitivity-aware caching framework. Specifically, we formalize the caching error through an analysis of the model output sensitivity to perturbations in the denoising inputs, i.e., the noisy latent and the timestep, and show that this sensitivity is a key predictor of caching error. Based on this analysis, we propose Sensitivity-Aware Caching (SenCache), a dynamic caching policy that adaptively selects caching timesteps on a per-sample basis. Our framework provides a theoretical basis for adaptive caching, explains why prior empirical heuristics can be partially effective, and extends them to a dynamic, sample-specific approach. Experiments on Wan 2.1, CogVideoX, and LTX-Video show that SenCache achieves better visual quality than existing caching methods under similar computational budgets.
Sources
- Building Normalizing Flows with Stochastic Interpolants
- Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
- $\Delta$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
- LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
- Explaining and Harnessing Adversarial Examples
- LTX-Video: Realtime Video Latent Diffusion
- Denoising Diffusion Probabilistic Models
- Imagen Video: High Definition Video Generation with Diffusion Models
- Video Diffusion Models
- Open-Sora Plan: Open-Source Large Video Generation Model
- Flow Matching for Generative Modeling
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
- MagCache: Fast Video Generation with Magnitude-Aware Cache
- Scalable Diffusion Models with Transformers
- FORA: Fast-Forward Caching in Diffusion Transformer Acceleration
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- Score-Based Generative Modeling through Stochastic Differential Equations
- Impact of QCD sum rules coupling constants on neutron stars structure
- Wan: Open and Advanced Large-Scale Video Generative Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models