Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching

arXiv:2412.18911 · cs.LG, cs.AI, cs.CV · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching".

Jane: The paper was written by Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu et al. from Shanghai Jiao Tong University and South China University of Technology and Beihang University and Huawei Technologies Ltd and Shanghai AI Lab and Hong Kong University of Science and Technology (Guangzhou).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we're digging into a paper that's been making the rounds on arXiv, and the title alone tells you it's doing some serious rethinking: "Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching."

Jane: And Tom, I have to say, when I first read that title, I thought, okay, another acceleration paper, we've seen a bunch. But this one actually flips a couple of assumptions on their head, and that's what got me excited.

Tom: Right, and for our listeners who might be new to this, let's set the stage. Diffusion models are the engines behind a lot of the image and video generation you see online. They work by starting with pure noise and gradually denoising it over many steps to reveal a picture. The problem is, those steps are computationally brutal.

Jane: Exactly. And the authors here, Zou, Zheng, Zhang, and their team from Shanghai Jiao Tong University and a few other places, they're looking at a specific trick called feature caching. The idea is that between two adjacent denoising steps, the features inside the model don't change that much. So why recompute everything from scratch?

Tom: So you cache the results from one step and reuse them in the next. That's the core idea. But here's where it gets interesting. There's a previous method called token-wise feature caching, or ToCa, that says, let's be smart about this. Some tokens in the image are more important than others, so let's compute those important ones fresh and only cache the unimportant ones.

Jane: And that sounds reasonable, right? But this paper asks two really pointed questions. First, do you actually need to compute those "important" tokens every single step? And second, are the tokens you're calling important actually important?

Tom: And the answers they found are genuinely counterintuitive. On the first question, they found that in the very first caching step, computing the important tokens doesn't really help because the error hasn't accumulated yet. You're just wasting compute. On the second question, they found that the so-called important tokens, selected by attention scores, sometimes perform worse than just picking tokens at random.

Jane: Which is wild. I mean, we've been conditioned to think that attention scores are the gold standard for importance in transformers. And here they're showing that random selection can beat it.

Tom: Yeah, and that's the hook for me. It challenges a pretty deep assumption in the field. We'll get into why they think that is and what they propose instead in a bit, but for now, just sit with that. The most sophisticated selection method might be worse than rolling dice.

Jane: And that's not just a theoretical curiosity. It has real practical implications for how we build these systems. But let's hold that thought, because next we need to talk about what they actually built to replace ToCa.

Tom: Good point, Jane. We've got the setup, now let's get into the meat of the method.

Summary: Jane: So we're back, and we're still on "Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching." Tom, let's get into what they actually propose, because the title says "dual feature caching," and that's the key innovation.

Tom: Right. So they call it DuCa for short. And the insight comes from looking at two existing caching strategies and realizing each one has a sweet spot. There's aggressive caching, where you skip almost all computation in a step and just reuse everything from the previous step. That gives you a huge speedup, but the error builds up fast if you do it for multiple steps in a row.

Jane: And then there's conservative caching, which is what ToCa does. It only caches some tokens, computes others, and that keeps the error in check but costs more compute per step.

Tom: Exactly. So DuCa's trick is to alternate between them. You do a fresh full computation step, then one aggressive caching step for maximum speed, then one conservative caching step to correct the drift, then aggressive again, and so on. It's like a rhythm. Aggressive to go fast, conservative to keep things aligned.

Jane: And that addresses their first question, right? They showed that computing the "important" tokens in the first caching step is unnecessary because the error hasn't accumulated. So you use aggressive caching there to get the speed. Then, once the error has had a chance to build up, you bring in the conservative step to fix it.

Tom: Precisely. And on the second question, about token selection, they ran a bunch of experiments comparing different scoring methods. Attention scores, K-norm, V-norm, similarity-based scores, and plain old random selection.

Jane: And the results in their Table I are pretty striking. On the FLUX model, random selection got an Image Reward of zero point nine eight zero six. The best attention-based method got zero point nine seven nine eight. So random actually beat it. And the similarity-based method, where you pick tokens that are most different from each other, got zero point nine eight one one, which was the best of all.

Tom: So the lesson they draw is that diversity matters more than importance. When you pick tokens that are similar to each other, you're getting redundant information. When you pick a diverse set, you're sampling the full picture. And random selection naturally gives you diversity without any extra computation.

Jane: And that's a huge practical win, because attention-score-based selection is incompatible with FlashAttention, which is this super-optimized attention implementation that everyone uses. So ToCa, by needing attention scores, was forced into a slower path. DuCa, by using random selection, works seamlessly with FlashAttention.

Tom: And that's why you see such a big difference in actual latency. In their experiments on FLUX, ToCa got a one point five nine times speedup in real time, but DuCa got two point nine three times. Same ballpark in FLOPs reduction, but nearly double the actual speed because it can use the fast hardware kernels.

Jane: So it's not just a theoretical improvement. It's a practical one that you feel when you're actually generating images. And the quality holds up too. They report an Image Reward of zero point nine eight nine six for DuCa at a three point four five times FLOPs reduction, compared to zero point nine seven three one for ToCa at three point three zero times. So DuCa is both faster and better.

Tom: And that's the summary in a nutshell. But we've only scratched the surface. Next we need to talk about how they validated this across different models and tasks, because they didn't just test it on one thing.

Jane: Good segue, Tom. Let's get into the experiments.

Improvements: Tom: Alright, we're back on "Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching." Jane, we've covered the core idea, but now I want to get into the breadth of their experiments, because they really put this thing through its paces.

Jane: They did. And that's one of the things I appreciate about this paper. They tested DuCa on text-to-image with FLUX and HiDream, text-to-video with OpenSora, and class-conditional image generation with DiT-XL/two on ImageNet. That's a solid spread.

Tom: And the results are consistently good. On OpenSora, which is a video model, they got a two point five zero times speedup with a VBench score of seventy-eight point eight three, which is actually higher than the original model's seventy-nine point four one minus a small drop. But compared to ToCa, which got a two point three six times speedup and a score of seventy-eight point five nine, DuCa is both faster and better.

Jane: And on ImageNet with DiT-XL/two they show that DuCa achieves a better FID than ToCa at similar acceleration ratios. For example, DuCa(b) gets an FID of two point eight four at a two point four eight times speedup, while ToCa(b) gets two point eight eight at two point three two times. So it's a clear win on both axes.

Tom: And let's talk about the ablation studies, because that's where they really dig into the design choices. They compared conservative-only, aggressive-only, and their alternating DuCa structure. Conservative-only gave an Image Reward of zero point nine seven eight nine but used a lot of FLOPs. Aggressive-only was cheap but dropped to zero point eight five three nine. DuCa hit zero point nine eight nine five with a middle ground in FLOPs.

Jane: That's a beautiful demonstration of the synergy. Neither extreme works well alone, but alternating between them gives you the best of both worlds. And they also tuned the cycle length and cache ratio. They found that a cycle length of five and a cache ratio of zero point nine gives the best trade-off.

Tom: And I think the deeper point here is about the philosophy of acceleration. A lot of methods try to be clever about where to skip computation. But this paper shows that sometimes the simplest approach, random selection, is the most robust. It's a humbling result.

Jane: It really is. And it also highlights the importance of thinking about the whole system, not just the algorithm. By being compatible with FlashAttention, DuCa gets a real-world speedup that's much larger than the FLOPs reduction alone would suggest. That's something a lot of papers overlook.

Tom: And that's the kind of insight that makes a paper actually useful in practice. You can have a beautiful algorithm that's slow because it fights the hardware. Or you can have a simple algorithm that works with the hardware and gets you the speed you need.

Jane: And the visual results in their paper back this up. In their Figure four you can see that DuCa's generated images are much closer to the original model's output than ToCa's, which sometimes has these weird grid-like artifacts. So it's not just a metric improvement, it's a visible quality improvement.

Tom: Right. And that's the kind of thing that matters when you're actually deploying these models in products. Users notice artifacts. They notice slowness. DuCa addresses both.

Jane: So we've covered the method, the experiments, and the improvements. Now let's wrap this up and think about what it all means.

Conclusion: Tom: And we're at the end of our time with "Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching." Jane, let's pull it all together for our listeners.

Jane: So the big takeaways are, first, that alternating between aggressive and conservative caching, rather than committing to one or the other, gives you the best balance of speed and quality. Second, that token selection based on importance scores is often worse than random selection, because what really matters is diversity.

Tom: And the practical impact is substantial. On FLUX, they get a three point four five times reduction in FLOPs with almost no quality loss, and a two point nine three times speedup in actual latency. On OpenSora, they get two point five zero times with a VBench score that's nearly identical to the original. That's not incremental. That's a real leap.

Jane: And I think the broader lesson here is about questioning assumptions. The field had settled on the idea that you need to identify important tokens to skip computation intelligently. This paper shows that maybe you don't. Maybe the simplest approach is the best.

Tom: And that's a valuable reminder for all of us. Sometimes the clever solution is actually the wrong one, and the humble solution is the right one. Random selection is about as humble as it gets.

Jane: And it also opens up new directions. If diversity is what matters, can we design even better selection strategies that are still compatible with FlashAttention? Maybe something that's not random but also doesn't require attention scores. That could be the next step.

Tom: Absolutely. And the authors also note that their method is training-free, which means you can apply it to existing models without any fine-tuning. That's huge for adoption. You just swap in DuCa and get the speedup.

Jane: So for anyone working on diffusion models, whether you're generating images, videos, or something else entirely, this paper is worth a read. It's a reminder that acceleration doesn't have to be complicated.

Tom: And with that, we're going to say goodbye to this paper. Thanks for joining us, and we'll be back soon with another one.

Jane: See you next time, everyone.

Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu, Zhengyi Shi, Conghui He, Xuming Hu, Linfeng Zhang

Shanghai Jiao Tong University · South China University of Technology · Beihang University · Huawei Technologies Ltd · Shanghai AI Lab · Hong Kong University of Science and Technology (Guangzhou)

cs.LG, cs.AI, cs.CV

Submitted: 2026-08-15

Updated: 2026-08-18

Code: https://github.com/black-forest-labs/flux

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 70/100

The gist: Paper Title: Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching Authors: Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu, Zhengyi Shi,

Key concepts

Dual Feature Caching (DuCa)
A method that alternates between two caching strategies: an aggressive one for maximum speed and a conservative one to correct error drift. This approach maintains performance while significantly reducing computational load.
Token-wise Feature Caching (ToCa)
A previous method where the model attempts to be smart by only caching unimportant tokens and recomputing the 'important' ones. The new research challenges this, finding that simpler methods can perform better than attention-based importance scoring.

Terminology

Summary

Paper Title: Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching

Core Contribution: The paper introduces DuCa (DUal CAching), a novel feature caching strategy for accelerating Diffusion Transformers (DiT) that performs aggressive caching and conservative caching iteratively, and replaces complex token selection methods with random selection.


Diffusion Transformers (DiT) have become the dominant methods in image and video generation but suffer substantial computational costs. Feature caching methods cache features from previous timesteps and reuse them in subsequent timesteps to skip computation. Among these, token-wise feature caching (ToCa) performs different caching ratios for different tokens, aiming to skip computation for unimportant tokens while computing important ones.

The paper poses two critical questions:

  1. Is it really necessary to compute the so-called 'important' tokens in each step?

  2. Are so-called important tokens really important?

The paper provides counterintuitive answers: consistently computing selected important tokens in all steps is not necessary, and the selection of important tokens is often ineffective, sometimes showing inferior performance than random selection.

The paper analyzes caching error (defined as the L2-norm distance between xt computed with and without feature caching) for aggressive caching versus conservative caching (ToCa). The authors find:

"We find that the caching error of the aggressive and conservative caching exhibit a close value at the first caching step (i.e., timestep 27 in Figure 2), suggesting that computing the important tokens in ToCa is not always necessary. Concretely, when the error from caching has not been accumulated yet, computing important tokens does not bring benefits and can be removed for better efficiency."

The paper evaluates token selection methods on FLUX.1-dev using Image Reward. The results show:

"Surprisingly, while the method of selecting tokens based on the highest attention score in ToCa is the best-performing approach among all standard-based selection methods, it still performs slightly worse than random selection."

The paper tests multiple selection criteria including Attention scores, K-norm, V-norm, and similarity-based methods. Key findings:

  • Selecting tokens with maximum similarity to base tokens achieved the worst performance among all selection methods

  • Selecting tokens with minimum similarity achieved the best performance

  • Random selection performs nearly as well as minimum similarity selection without additional computation

The authors conclude:

"We argue that token selection should focus on their duplication instead of their importance. Thus, random selection can be a good choice since randomness guarantees choosing tokens with different semantic information and leads to no additional computation or memory overhead."

Aggressive Caching: Directly replaces the output of the DiT block with cached features from previous layers, skipping all computations of the first l layers. Formulated as: gl:= C[l], where C[l] denotes the cached feature at the lth layer. The computation becomes: G(x)t:= gl+1 ◦... ◦ gL (C[l]). The paper sets l = L−1 in most experiments, skipping almost all layers in the caching step.

Conservative Caching: Caches and reuses only features without the residual connection, allowing computation in some layers to correct cached features. Following ToCa, an importance score S is calculated for each token xi, dividing tokens into ICache (cached) and ICompute (computed). The computation for token xi on layer l is: γi f l (xi) + (1 − γi)C[l](xi), where γi = 0 for cached tokens and γi = 1 for computed tokens.

DuCa alternates between aggressive and conservative caching:

"In each caching cycle, DuCa initializes the cache by performing the full computation in the fresh timestep. Then, it performs the one-step aggressive caching for high-ratio acceleration, followed by one-step conservative caching to fix the cache error."

Specifically:

We arrange token-wise conservative caching steps on odd-numbered steps following a fresh step and aggressive caching steps on even-numbered steps.

DuCa uses random scores for token selection in conservative caching steps:

The proposed method selects tokens for DuCa by setting random scores for all tokens to diversity. The tokens with the largest random scores are selected as cached tokens.

This approach provides compatibility with FlashAttention:

"The attention map-based selection strategy is incompatible with another commonly used attention-optimization method FlashAttention and memory-efficient attention, which can not provide the attention scores but are capable of reducing the memory costs of attention from O(N2) to O(N) and significant accelerating the attention layers."

  • DuCa (N=5, R=90%) achieves 3.45× FLOPs speedup with Image Reward of 0.9896, nearly unchanged from the original model's 0.9898

  • ToCa achieves only 3.30× FLOPs speedup with Image Reward of 0.9731 (a decrease of 0.0167)

  • DuCa reduces quality loss to 1.2% of that observed with ToCa

  • Latency speedup: DuCa achieves 2.93× versus ToCa's 1.59× due to FlashAttention compatibility

  • Even at 4.05× acceleration, DuCa maintains near-lossless GenEval performance

  • DuCa (N=4, R=90%) achieves 3.18× latency speedup and 3.54× FLOPs speedup with Image Reward of 1.090

  • Outperforms ToCa (1.56× latency, 3.06× FLOPs, Image Reward 1.111) and PAB

  • DuCa achieves 2.50× speedup with VBench score of 78.83%

  • ToCa achieves 2.36× speedup with VBench score of 78.59%

  • Original model: 79.41% VBench

  • DuCa reduces quality loss by 29.3% compared to ToCa

  • Generation time: 19.16s per video versus 29.03s for ToCa (nearly 10s faster)

  • DuCa(b) achieves similar FID-50k as ToCa(b) but with higher acceleration (2.48× vs 2.32×)

  • At 2.7× acceleration, DuCa achieves FID-50k of 3.00 versus ToCa's 3.31

The paper demonstrates that diversity matters more than importance in token selection. Similarity-based experiments confirm that maximum similarity selection performs worst, while minimum similarity performs best.

  • Conservative-only: Good quality (Image Reward 0.9789) but high computation (1263.22 TFLOPs)

  • Aggressive-only: Low computation (1042.12 TFLOPs) but poor quality (Image Reward 0.8539)

  • DuCa: Balanced (1190.24 TFLOPs, Image Reward 0.9895) — better quality than both alternatives

  • Cycle length N=5 provides the best trade-off between computation and quality

  • Cache ratio R=0.9 eliminates redundant computation without compromising fidelity

  • Lowering R below 0.7 significantly harms quality

The paper concludes:

"This paper first gives an in-depth study of the caching error on aggressive and conservative methods and then introduces dual caching DuCa, which performs aggressive and conservative alternatively to leverage their advantages in different caching steps. Moreover, we have discovered the value of random selection and further explored its implications, revealing that diversity plays a more critical role in token selection than previously thought."

The authors emphasize that random selection not only demonstrates excellent performance but also seamlessly integrates with efficient attention methods like FlashAttention, making feature caching a practical and viable approach in real-world applications.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Replace single-strategy feature caching with an alternating aggressive/conservative caching scheme.

Implementation:

  • After a full computation step, perform one aggressive caching step (skip all layers) followed by one conservative caching step (compute only selected tokens)

  • Use cycle length N=5 with 90% cache ratio as optimal hyperparameters

  • This reduces FLOPs by 3.45× on FLUX.1-dev while maintaining Image Reward within 0.02% of the original model

What the improved system can do:

  • Generate 1000×1000 images in 5.88 seconds instead of 17.20 seconds (2.93× latency speedup)

  • Maintain generation quality nearly identical to the uncompressed model

  • Reduce quality loss to 1.2% of what previous token-wise caching (ToCa) achieves

Sources

Related papers