From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion
Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang
Fudan University · The Chinese University of Hong Kong
cs.AI, cs.CV, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/black-forest-labs/flux
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead.
Terminology
Summary
Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which are significantly misaligned with final generation quality. This discrepancy stems from the non-uniform propagation and accumulation of errors along the denoising trajectory. To address this, the paper proposes Global-Impact Cache (GCache). The method establishes a rigorous theoretical characterization of the error propagation upper bound. Recognizing that this bound can be overly conservative for complex, highly non-convex diffusion models, the authors reparameterize the propagation exponent with a Bernstein form and reformulate cache policy search as a bilevel optimization problem. GCache identifies an optimal reuse policy in the inner objective while aligning the error-weighting function with generation quality loss in the outer objective. Extensive experiments demonstrate that GCache consistently outperforms prior caching strategies on both video and image generation. Notably, on the state-of-the-art Wan2.1 video diffusion model, GCache maintains a 2.17× speedup while significantly enhancing generation quality, reducing LPIPS from 0.1095 to 0.0316.
Diffusion models have emerged as the dominant paradigm in visual generation, delivering high-fidelity and diverse outputs across various modalities, including images and videos. Despite these successes, diffusion models remain burdened by substantial inference overhead, stemming from the iterative nature of solving the underlying Ordinary Differential Equations (ODEs), which necessitates numerous model evaluations for a single sample. Various acceleration mechanisms have been explored, including model-centric compression, advanced sampling solvers, and cache-based mechanisms.
Cache-based methods provide a practical approach to speeding up diffusion inference. Unlike model-centric techniques, caching avoids intensive retraining or distillation of model parameters and remains orthogonal to advanced sampling solvers, making it a highly lightweight solution for denoising acceleration. The primary objective is to establish an inference-time policy that identifies redundant intermediate residuals across adjacent timesteps to avoid repeated computations. While initial approaches relied on uniform reuse schedules, they lacked the flexibility to adapt to varying residual dynamics across different timesteps. Recent works have shifted toward non-uniform strategies, employing local similarity metrics to quantify the mismatch between target and cached residuals, triggering reuse only when the discrepancy is sufficiently small.
The paper identifies that existing approaches remain limited by their reliance on local similarity metrics, such as the relative l1 distance. Empirical evidence suggests that such locally-based strategies fail to accurately capture how individual cache reuse decisions affect final generation quality. Large local discrepancies may correspond to only minor perceptual degradation: pronounced relative l1 spikes occur around step 19 and increase sharply in the final denoising stages, yet result in only marginal increases in LPIPS. This mismatch reveals that local discrepancy alone is an unreliable proxy for global generation impact, indicating that effective cache reuse policies must account for both local reuse errors and their cumulative impact on final generation quality.
The paper first conducts controlled perturbation experiments on video generation, injecting random noise of a fixed magnitude into intermediate feature maps at different denoising timesteps and tracking the log-scaled deviation ∥xt − x̂t∥1 over subsequent denoising steps. The results show that perturbations introduced earlier are progressively amplified and lead to substantially larger final deviations than those injected later. These findings reveal a cumulative and time-dependent error propagation effect, suggesting that an effective cache reuse policy should account for not only the local cache reuse error but also the error propagation inherent in the denoising process.
The paper establishes formal theoretical results under standard regularity conditions (Assumption 3.1), including bounded approximation error, Lipschitz continuity, and bounded velocity dynamics. Theorem 3.2 provides a global error bound under cache reuse, decomposing the final error into three components: (i) the propagated error originating from cache reuse, (ii) the network approximation error stemming from the learned velocity, and (iii) the discretization error inherent to the ODE solver. The cache reuse term is scaled by an exponential factor, indicating that local errors introduced by cache reuse are amplified by the system dynamics as the trajectory evolves.
Theorem 3.3 establishes an error propagation bound for single-step cache reuse, showing that the cumulative error at the final timestep is bounded by ∥ϵcti+1∥1 e wti+1, where the propagation exponent wti+1 is defined as wti+1 = ln(hti+1 Lout) + Σ n=1 i-1 htn+1 Ltn+1. Theorem 3.4 extends this to multi-step cache reuse, providing a bound that sums over all reuse events.
These bounds suggest that an optimal cache reuse policy can be identified by minimizing the upper bound of the total propagation error. The policy optimized via this theoretical bound consistently achieves lower reconstruction error under a fixed computation budget compared to the state-of-the-art ERTACache, demonstrating the bound's ability to capture long-term dependencies of errors.
While the analytical bound serves as a robust guide, its worst-case nature inherently introduces a pessimistic bias, as it does not fully account for the high non-convexity and intrinsic error-resilience of diffusion models. The paper proposes Global-Impact Cache (GCache), which parameterizes the propagation exponent wt using the d-th degree Bernstein polynomials: w(t; s) = Σ ν=0 d s ν+1 (d choose ν) t ν (1-t) d-ν, where s denotes a vector of learnable parameters.
The policy search is formulated as a bilevel optimization problem:
-
Outer objective: s = arg min L(m⋆(s)), optimizing parameters s to minimize a generation quality metric (e.g., LPIPS)
-
Inner objective: m⋆(s) = arg min Σ ∥ϵcti+1∥1 e w(ti+1;s), finding the optimal cache reuse policy that minimizes the current error bound
The inner optimization is solved via Dynamic Programming with O(KN2) complexity, treating the problem as a constrained shortest path problem on a directed acyclic graph. The outer optimization uses Bayesian Optimization with a Gaussian Process surrogate and Lower Confidence Bound acquisition function, since evaluating the outer objective requires a full inference pass and is a high-cost black-box function.
The paper evaluates on four representative DiT-based diffusion models: three video diffusion backbones (Open-Sora 1.2, CogVideoX, and Wan 2.1) and one text-to-image model (Flux-dev 1.0). Video generation is evaluated using the official 946 prompts from VBench, and image generation using the official 30K prompts from COCO. Metrics include VBench, LPIPS, PSNR, and SSIM.
Across all evaluated settings, GCache consistently achieves state-of-the-art performance, delivering both faster inference and higher generation quality than existing cache-based acceleration methods. Key results include:
-
On Open-Sora 1.2: GCache-slow (K=18) achieves 1.56× speedup with LPIPS of 0.1363, compared to ERTACache-slow's 1.55× speedup with LPIPS of 0.1659
-
On CogVideoX-2B: GCache-slow (K=31) achieves 1.62× speedup with LPIPS of 0.0721, compared to ERTACache-slow's 1.62× speedup with LPIPS of 0.1012
-
On Wan2.1-1.3B: GCache-slow (K=24) achieves 2.17× speedup with LPIPS of 0.0316, compared to ERTACache's 2.17× speedup with LPIPS of 0.1095; GCache-fast (K=16) achieves 3.01× speedup with LPIPS of 0.0828
-
On Flux-dev 1.0: GCache-fast (K=10) achieves 2.87× speedup with LPIPS of 0.1825, compared to ERTACache's 2.87× speedup with LPIPS of 0.2658
Qualitative comparisons on Wan 2.1 show that ERTACache suffers from motion and object misalignment under aggressive cache reuse, including misaligned motion dynamics and object-level deviations from original outputs. In contrast, GCache preserves consistent motion patterns and object semantics. On Flux-dev 1.0, ERTACache exhibits clear semantic and spatial misalignment, such as generating four smoking stacks when the prompt specifies two, while GCache maintains semantic correctness and spatial coherence.
-
Polynomial degree d: The best performance is achieved at d=3, yielding the lowest LPIPS and highest SSIM and PSNR
-
Outer objective: The hybrid LPIPS+SSIM objective achieves the best overall performance across all metrics, with the lowest LPIPS (0.0721) and highest PSNR (29.14)
-
Robustness to prompt distribution shifts: Performance variance across different training sources (Static, Dynamic, Mixed) is remarkably marginal, demonstrating robust generalization
-
Generalization across resolutions: GCache-fast policy optimized at 1024×1024 transfers effectively to 512 and 256 resolutions without re-tuning, consistently outperforming ERTACache
-
Policy effectiveness: GCache outperforms both ERTACache and ERTACache* (without error rectification) under identical budgets, demonstrating that a well-optimized policy can be more powerful than a sub-optimal policy supplemented by error correction
-
Pre-computed error proxy validation: The error profiles of original and cache-reused trajectories are remarkably aligned, justifying the use of pre-computed errors as a reliable proxy
The paper identifies the misalignment between local reuse discrepancies and global generation error, establishing a formal theoretical characterization of error propagation dynamics in cache-based acceleration. The authors demonstrate that policies optimized strictly via conservative analytical bounds are often sub-optimal in practice. To bridge this gap, they introduce Global-Impact Cache (GCache), a framework that reformulates the policy search as a bilevel optimization problem, effectively reconciling theoretical error control with empirical perceptual quality. Extensive evaluations across various image and video backbones show that GCache consistently outperforms prior caching strategies.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
1. Implement Global-Impact Cache (GCache) for diffusion model inference acceleration
-
I can integrate the bilevel optimization framework into existing diffusion model pipelines to replace local similarity-based caching policies
-
The improved system can achieve 2–3× faster inference on video and image generation while improving output quality (e.g., reducing LPIPS from 0.1095 to 0.0316 on Wan2.1) rather than trading quality for speed
2. Add error-propagation-aware caching decisions
-
I can use the theoretical error propagation bounds (Theorem 3.2–3.4) to weight cache reuse errors by their cumulative impact on the final output, not just local discrepancy
-
The improved system can avoid quality degradation in early denoising steps where errors amplify, leading to more stable motion and object consistency in video generation
3. Learn optimal cache reuse policies via Bayesian optimization
-
I can employ the outer-loop Bayesian optimization with Gaussian Process surrogate to tune the Bernstein-polynomial error-weighting parameters for each specific model and task
-
The improved system can automatically adapt its caching policy to different diffusion backbones (e.g., Open-Sora, CogVideoX, Wan2.1, Flux-dev) without manual tuning, maintaining high quality across diverse generation tasks
4. Use pre-computed error profiles as a reliable proxy for policy search
-
I can leverage the validated alignment between original and cache-reused trajectory error profiles to pre-compute error statistics and avoid expensive online evaluations
-
The improved system can perform fast, one-time policy optimization per model and then apply it to any prompt, reducing inference overhead further
5. Enable resolution- and distribution-robust acceleration
-
I can apply the GCache policy trained at one resolution (e.g., 1024×1024) to other resolutions (512, 256) without re-tuning, and maintain performance under prompt distribution shifts
-
The improved system can be deployed across varied input sizes and content types while retaining both speed and quality, making it practical for real-world applications
6. Integrate hybrid quality objectives for policy optimization
-
I can use the hybrid LPIPS+SSIM objective (which the paper shows yields the best overall metrics) to guide cache policy search, balancing perceptual and structural fidelity
-
The improved system can produce outputs that are not only fast but also more perceptually aligned with human judgment, reducing visible artifacts in both images and videos
7. Replace sub-optimal error-correction mechanisms with better policies
-
I can show that a well-optimized cache policy (GCache) outperforms even an error-corrected baseline (ERTACache*) under the same compute budget, so I can remove unnecessary error-rectification modules
-
The improved system becomes simpler and more efficient, avoiding extra computation while achieving higher generation quality
8. Provide a theoretical guarantee for cache reuse decisions
-
I can use the formal error bounds to set safe reuse thresholds, ensuring that any caching decision does not exceed a provable worst-case deviation from the full computation
-
The improved system can offer reliability guarantees for safety-critical or high-stakes generation tasks (e.g., medical imaging, autonomous driving simulation) where uncontrolled error is unacceptable
9. Extend the framework to other iterative generative models
-
I can generalize the bilevel cache optimization to other ODE-solver-based generative models (e.g., flow matching, consistency models) that share similar error propagation dynamics
-
The improved system can accelerate a broader class of generative AI models beyond diffusion, increasing deployment efficiency across the AI ecosystem
10. Enable real-time interactive generation
-
By combining GCache’s speedup with its quality preservation, I can build systems that generate high-fidelity images or videos in near-real-time (e.g., 3× speedup on Flux-dev with only 0.18 LPIPS)
-
The improved system can support interactive creative tools, live video editing, or on-device generation where both latency and quality are critical
Abstract
Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation quality. This discrepancy stems from the non-uniform propagation and accumulation of errors along the denoising trajectory. To address this, we propose Global-Impact Cache (GCache). We first establish a rigorous theoretical characterization of the error propagation upper bound. Recognizing that this bound can be overly conservative for complex, highly non-convex diffusion models, we further reparameterize the propagation exponent with a Bernstein form and reformulate cache policy search as a bilevel optimization problem. In detail, GCache identifies an optimal reuse policy in the inner objective while aligning the error-weighting function with generation quality loss in the outer objective. This framework effectively reconciles theoretical rigor with empirical performance, learning to prioritize computation where it most impacts visual fidelity. Extensive experiments demonstrate that GCache consistently outperforms prior caching strategies on both video and image generation. Notably, on the state-of-the-art Wan2.1 video diffusion model, GCache maintains a 2.17x speedup while significantly enhancing generation quality, reducing LPIPS from 0.1095 to 0.0316.
Sources
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- $\Delta$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
- FORA: Fast-Forward Caching in Diffusion Transformer Acceleration
- Open-Sora: Democratizing Efficient Video Production for All
- Wan: Open and Advanced Large-Scale Video Generative Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection