RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models
summary
The gist
Putting GEMM Sparsity in the Right Place for Diffusion Models" addresses the inference cost bottleneck of Diffusion Transformers (DiT) in image generation.
In short
The episode discusses 'RT-Lynx,' a paper detailing how to apply sparsity in diffusion models. The hosts conclude that applying sparsity to activations, rather than static weights, is crucial for maintaining image quality and improving speed. RT-Lynx achieves significant performance gains while preserving fidelity across various architectures through specialized techniques.
Key concepts
- GEMM
- General Matrix Multiplication (GEMM) is the core computational operation in neural networks. The efficiency of this calculation determines how fast a model runs. The paper focuses on making these operations sparse to reduce the amount of data that needs to be processed.
- Sparsity (Activation vs. Weight)
- Sparsity refers to having many zero values in a matrix or data flow. The hosts discuss that pruning weights is static, but pruning activations—the dynamic data flowing through the network—is a much harder and more effective engineering challenge for diffusion models.
- RT-Lynx
- This is the framework developed by Alibaba Group to implement sparsity efficiently. It fuses the sparsification process directly into custom CUDA kernels, allowing it to run on-the-fly during inference. This reduces overhead and enables real-time image generation speedups.
- LoRA (Low-Rank Adapter)
- LoRA is a small adapter trained to compensate for information lost when making a model sparse. It adds a low-rank correction term to the sparse computation, helping restore fine details like textures and edges that might have been pruned.
Terminology used across episodes
This episode discusses
- RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models · Paper Radio
- Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
- Perception Encoder: The best visual embeddings are not at the output of the network
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- -DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
- SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
- Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
- Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
- SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
- E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
- Flow Matching for Generative Modeling
- Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
- ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
- Training-Free Activation Sparsity in Large Language Models
- From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
- La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation
- Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models
The paper
RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models · Read on arXiv
Xing Cong, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Chenhao Xie
Alibaba Group
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models".
Jane: The paper was written by Xing Cong, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu et al. from Alibaba Group.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we're digging into a fresh arXiv paper that's got a great title: "RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models." Tom here, and I've got Jane with me, plus our regulars Lu and Meng. Jane, what do you make of that title?
Jane: I love it, Tom. "Putting GEMM sparsity in the right place" — that's the whole thesis right there. GEMM is just matrix multiplication, the bread and butter of neural networks. And these authors from Alibaba Group are saying we've been applying sparsity to the wrong part of the model.
Lu: Exactly, Jane. For years, the field has focused on sparsifying weights — pruning away the less important numbers in the model's parameters. But this team looked at diffusion models, the ones generating images, and found that the activations — the data flowing through the network — are naturally much sparser.
Meng: And that's a big deal for someone like me who cares about actually running these models. Sparse weights are static, you can pre-process them. Sparse activations are dynamic, they change every single inference. That's a much harder engineering problem.
Tom: Right, Meng, and that's why the title says "the right place." The authors show that when you force a two:four sparsity pattern — keeping only two out of every four values — on the weights of a diffusion model, the image quality just collapses. But doing the same to activations barely hurts at all.
Jane: They even have this beautiful figure in the paper showing a capybara wearing a suit holding a sign. The weight-sparse versions produce absolute garbage — like, unrecognizable mush. But the activation-sparse version looks nearly identical to the original.
Lu: The intuition is fascinating. In large language models, researchers found this thing called superposition — each token activates only a small subset of neurons. The authors show the same thing happens in diffusion transformers. The activations are concentrated near zero, with only maybe five to ten percent of neurons really firing.
Meng: So you're pruning away values that are already close to zero. Of course that hurts less than chopping out weights that might be critical.
Tom: And that's the core insight that flips the whole field on its head. The authors are from Alibaba, and they've built an entire system around this — the RT-Lynx framework. We're going to dig into how they actually make this work in practice.
Jane: I'm curious about the engineering side too, because dynamic sparsity sounds like a nightmare to implement efficiently. But first, let's talk about what they actually found when they compared weight sparsity to activation sparsity across different models.
Paper Summary: Tom: So we're back with "RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models." Jane, you were saying you wanted to dig into the actual results.
Jane: Yeah, Tom. The numbers are pretty stark. On Qwen-Image, a popular open-source diffusion model, they measured FID — that's a metric for how realistic generated images look. The full dense model scores about twenty-two. If you apply naive weight sparsity, it jumps to over fifty-one. That's a catastrophic drop in quality.
Lu: But here's the kicker — naive activation sparsity only pushes it to about thirty-six. Still not great, but way better. And then when they add their compensation techniques, they actually beat the original model, getting FID down to twenty-one point two five. That's not just recovering the loss, that's improving on the baseline.
Meng: Wait, Lu, how do you beat the original? That seems suspicious.
Lu: It's not magic. They're using a LoRA branch — a low-rank adapter — that's trained to compensate for exactly what the sparsification removes. And because the sparse activations are actually acting as a form of regularization, the model ends up generating cleaner images in some cases.
Tom: And they tested this across four different model configurations — Qwen-Image, Qwen-Image-two thousand five hundred twelve FLUX.one-dev, and Z-Image. The pattern holds everywhere. Weight sparsity wrecks the models, activation sparsity with their fixes keeps quality intact.
Jane: The visual examples in the paper are really compelling. There's this one prompt about a rainy woods scene, and the weight-sparse version produces this weird, distorted mess. The activation-sparse version with RT-Lynx looks almost indistinguishable from the original.
Meng: So the quality story is solid. But what about speed? Because if you're doing dynamic sparsification at inference time, that's extra compute. You have to figure out which activations to keep, then reformat them, then do the sparse matrix multiply.
Tom: That's exactly the challenge they tackle, Meng. And this is where the engineering gets interesting. They built custom CUDA kernels that fuse the sparsification step directly into the matrix multiplication. Instead of doing it as a separate pass, it happens on the fly inside the kernel.
Lu: The overhead drops from something like forty percent of runtime down to under ten percent. And they're getting up to one point eight eight times speedup on the sparse GEMM itself, which translates to about one point five five times on the linear layers overall.
Jane: And end-to-end, that's about one point two times faster for generating an image. Which might not sound huge, but when you're running thousands of generations, that adds up.
Meng: I want to know more about how they actually implement this. Fusing sparsification into the kernel is clever, but there's got to be more to it.
Jane: That's our next segment — the technical improvements they made. Let's get into the weeds.
Improvements: Tom: We're back with "RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models." Jane just teased the technical side. Meng, you had questions.
Meng: Yeah, Tom. The paper describes three main improvements. First, there's norm compensation. When you prune activations, you're reducing the overall magnitude of the data. So they rescale the remaining values to match the original norm. It's a simple trick but it recovers a lot of the lost signal.
Lu: And it's elegant because it costs almost nothing computationally. Just a scaling factor per token. But it closes a big chunk of the quality gap — FID goes from thirty-five point eight five down to twenty-five point two eight on Qwen-Image just from that one change.
Jane: Then there's the LoRA branch. They train a low-rank adapter that takes the original dense activations and produces a correction term. So the final output is the sparse computation plus this low-rank residual. The LoRA is tiny — rank sixty-four — so it adds maybe ten percent overhead, but it recovers the fine details like hair textures and edges.
Meng: And the third piece is selective layer skipping. For some models, especially the single-stream architectures like FLUX and Z-Image, certain layers are just too sensitive to sparsification. So they skip those layers entirely — leave them dense.
Lu: That's a pragmatic engineering choice. They measured the relative error each layer type introduces, and found that the mixed text-image layers in single-stream models are the most fragile. So they skip the output projection and one of the MLP layers in those models.
Tom: And then there's the CUDA kernel work, which is honestly the most impressive part to me. They fuse the entire sparsification pipeline — pattern determination, top-k selection, compression — into a single execution path at the register level.
Meng: Right, Tom. The traditional approach would be: sparsify the activations, write them to memory, then read them back for the sparse GEMM. That round-trip is expensive. RT-Lynx keeps everything on-chip, generates the sparse format directly in registers, and feeds it straight into the sparse tensor cores.
Jane: They also interleave the sparse computation with the dense LoRA branch. So while the sparse tensor cores are doing their thing, the dense tensor cores are computing the LoRA correction. Then the results are accumulated on-chip. No intermediate memory writes.
Lu: The result is that the online sparsification overhead drops to as low as two point two eight percent of total execution time. That's the number that makes this practical. If the overhead were still forty percent, the whole approach would be dead in the water.
Meng: And they show it composes with other acceleration techniques. You can stack RT-Lynx on top of quantization, on top of step distillation, on top of caching. Each one gives you an additional speedup without breaking the others.
Tom: So we've got quality preserved, speed improved, and compatibility with existing optimizations. That's a pretty complete package. Jane, what do you think the bigger picture is here?
Jane: I think this is one of those papers that makes you reconsider assumptions the whole field has been operating under. Let's wrap up with that thought.
Conclusion: Tom: Alright, we're wrapping up our discussion of "RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models." Jane, what's the big takeaway for our listeners?
Jane: The big takeaway is that sparsity isn't a one-size-fits-all tool. The field spent years perfecting weight sparsification for language models, and when diffusion models came along, everyone assumed the same approach would work. This paper shows that assumption was wrong.
Lu: And it's not just that activation sparsity works better — it's that the structure of the data makes it fundamentally better suited. Diffusion models process thousands of tokens per image, and those tokens have this natural sparsity that weights just don't have.
Meng: From my side, the engineering is what makes it real. A theoretical insight about sparsity doesn't help anyone if you can't actually run it faster. The fact that they built custom kernels that get real speedups on real hardware — that's what moves the needle.
Tom: And the numbers speak for themselves. Up to one point eight eight times on the sparse GEMM, one point five five times on linear layers, and about one point two times end-to-end. Plus it stacks with quantization and distillation for even bigger gains.
Jane: The authors also showed it works across multiple architectures — Qwen-Image, FLUX, Z-Image. So this isn't a one-off hack. It's a general technique.
Lu: I think the biggest implication is for deployment. Diffusion models are expensive to run, which limits who can use them. If you can get a twenty percent speedup essentially for free, that lowers the barrier for smaller teams and real-time applications.
Meng: And the fact that it's plug-and-play with existing acceleration methods means people don't have to choose between RT-Lynx and other optimizations. They can have both.
Tom: So we're saying goodbye to "RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models" — a paper that might just change how we think about sparsity in generative models. Thanks to Lu and Meng for joining us today.
Jane: And thanks to all our listeners. Next episode, we're looking at a paper on efficient attention mechanisms for video generation. You won't want to miss it.
Tom: Until then, keep exploring. This is Tom and Jane, signing off.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language