CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts

summary

Video file (mp4)

In short

CodeQuant addresses the challenge of compressing large Mixture-of-Experts AI models by tackling precision loss caused by outliers. The paper introduces a three-stage framework—outlier smoothing, adaptive weight clustering, and weight reordering—achieving up to 4.15x speedup and massive memory savings while maintaining high accuracy for practical edge device deployment.

Key concepts

Mixture-of-Experts (MoE) Models
These models operate like a team of specialists rather than one giant brain. A router directs incoming data to the most appropriate 'expert' module, allowing developers to build massive, powerful AI models that remain efficient and fast for use.
Quantization
This is the process of making large AI models smaller by representing numbers using fewer bits. While necessary for deployment on consumer hardware, this process can cause significant loss of precision, especially when dealing with extreme data values called outliers.
Outliers in AI
In data, outliers are values that are significantly larger or smaller than the rest of the dataset. When compressing a model using quantization, these few extreme outliers can disrupt the entire process and severely degrade the model's overall accuracy.

Terminology used across episodes

This episode discusses

The paper

CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts · Read on arXiv

Xiangyang Yin, Xingyu Liu, Tianhua Xia, Bo Bao, Vithursan Thangarasa, Valavan Manohararajah, Eric Sather, Sai Qian Zhang

New York University · Cerebras Systems Inc.

Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increasingly central to large-scale language modeling. Under post-training quantization (PTQ), these outliers induce substantial quantization errors, leading to severe accuracy degradation. While recent rotation-based smoothing techniques alleviate the problem by redistributing outlier magnitudes, residual errors remain and continue to impede reliable low-precision deployment. In this work, we tackle this challenge by introducing CodeQuant, a unified quantization-and-clustering scheme that contains smoothing activation outliers via learnable rotation and absorbing weight outliers into fine-tuned cluster centroids for MoE. This design reduces the influence of extreme values by fitting them within cluster centroids, thereby lowering quantization error while maintaining expressive capacity. Coupled with a dedicated kernel design for GPU and CPU, CodeQuant achieves up to 4.15 times speedup while delivering significantly higher accuracy than state-of-the-art quantization approaches across diverse MoE models. Our results highlight CodeQuant as a promising direction for efficient and accurate deployment of MoE-based large language models under low-precision constraints. Our code is available at https://github.com/SAI-Lab-NYU/CodeQuant.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts".

Jane: The paper was written by Xiangyang Yin, Xingyu Liu, Tianhua Xia, Bo Bao, Vithursan Thangarasa et al. from New York University and Cerebras Systems Inc..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone! Today we’re diving into a fresh arXiv paper that’s got me genuinely pumped. It’s called "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts." Jane, that title is a mouthful, but the problem it solves is something we’ve talked about a lot on this channel.

Jane: Absolutely, Tom. And let me just say, for our listeners who aren’t deep in the weeds of machine learning, this paper is about making these massive AI models smaller and faster without breaking their brains. We’re talking about Mixture-of-Experts models, which are like a team of specialists. Instead of one giant brain doing everything, you have a router that sends each question to the right expert. That’s how you get models like Mixtral or Qwen that are huge but still fast.

Tom: Right, and the catch is, when you try to compress these models to run on your phone or a regular server, you run into these things called "outliers." Jane, you want to explain what those are for the folks at home?

Jane: Sure. Imagine you’re trying to take a photo of a room, but one lightbulb is blindingly bright. If you lower the exposure to handle that one bulb, the whole room goes dark. Outliers in AI are like that blinding lightbulb—a few values in the data that are way bigger than everything else. When you try to compress the model, those outliers mess up the whole process.

Tom: And that’s where CodeQuant comes in. The team, from NYU and Cerebras Systems, has a clever way of smoothing out those outliers and then using a technique called clustering to pack the weights more efficiently. We’re going to get into the nitty-gritty in a second, but the headline is that they’re getting up to four point one five times speedup on some hardware while keeping accuracy that blows the other methods out of the water.

Jane: It’s a big deal because it’s not just about making things smaller; it’s about making them usable. We’re talking about running these powerful models on edge devices, which could change how we interact with technology daily.

Tom: Exactly. And I want to bring in our resident expert, Lu, to give us the high-level view. Lu, what’s the big picture here?

Lu: Thanks, Tom. The big picture is that we’re hitting a wall with how much compute and memory these models need. CodeQuant is a smart workaround. Instead of fighting the outliers, they’re essentially redistributing the light in that room so you can take a better photo. They rotate the data to make the outliers less extreme, and then they cluster the weights into groups that are easier to compress. It’s a two-pronged attack that’s proving to be much more effective than the one-size-fits-all quantization we’ve seen before.

Tom: And we’re just getting started. Stick around because we’re going to break down exactly how they do this magic and what it means for the future of AI.

Summary: Jane: Welcome back. We’re still on "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts." Tom, we gave the elevator pitch, but let’s get into the actual summary of what this paper does.

Tom: Right. So, the core problem is that when you quantize a model—that’s the process of using fewer bits to represent numbers—you lose precision. And with outliers, you lose a ton of precision. The authors, led by Xiangyang Yin and Xingyu Liu, propose a three-stage framework to fix this. First, they use something called "Activation-Oriented Outlier Smoothing." That’s their fancy way of saying they apply a learned rotation to the data to make the outliers less extreme before quantization.

Jane: And that’s the "smoothing" part. But they don’t stop there. They also introduce "Adaptive Weight Clustering." Instead of just rounding numbers to the nearest value, they group similar weights together and find the best representative value for each group. It’s like organizing a messy closet by grouping all your t-shirts together and then picking one perfect t-shirt to represent them all.

Lu: And that’s a great analogy, Jane. But the clever part is that they don’t just pick the groups randomly. They fine-tune the centroids—the representative values—to minimize the actual error in the output. They’re not just looking at the weights in isolation; they’re looking at how the weights interact with the activations to produce the final result. That’s what makes their clustering so much better than the standard K-means approach.

Tom: And there’s a third piece called "Permutation-Invariant Outlier Grouping." This one is a bit more technical, but essentially, they reorder the weights before clustering. This makes sure that the extreme outliers are spread out across different groups instead of all ending up in one group and ruining that group’s accuracy.

Jane: So they smooth, they cluster, and they reorder. And the results are pretty stunning. On models like Mixtral eight times 7B, they’re getting perplexity scores—that’s a measure of how confused the model is—down to four point six five, which is incredibly close to the original model’s four point zero one. Other methods were blowing up to sixteen or even seventy-seven.

Lu: That’s the key takeaway. The accuracy is almost preserved, even at four-bit precision. And they show this across a range of models, from smaller ones like Phi-mini-MoE to the massive Qwen3-30B-A3B. The consistency is what’s impressive.

Tom: And we haven’t even talked about the speed yet. That’s coming up in the next segment, so stay tuned.

Improvements: Jane: Welcome back to the show. We’re deep into "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts." Tom, we’ve covered the problem and the method. Now let’s talk about the actual improvements this paper brings to the table.

Tom: Absolutely, Jane. And I want to bring in Meng, our engineer, because this is where the rubber meets the road. The paper doesn’t just stop at the algorithm; they actually built a custom kernel to make this thing run fast on real hardware.

Meng: Yeah, and that’s the part that gets me excited. A lot of papers propose clever algorithms, but they’re too slow to actually use. CodeQuant is different. They designed a lookup-table-based kernel. Instead of doing a bunch of multiplications, they precompute the results and just look them up. It’s like having a multiplication table instead of calculating seven times eight every time you need it.

Tom: And that’s where the speedup comes from. They’re reporting up to two point six three times speedup on an A100 GPU compared to the standard BF16 model. And on a CPU, they’re seeing up to four point one five times speedup. That’s a huge deal for deployment.

Jane: And it’s not just about speed. They’re also saving a ton of memory. On the Qwen3-30B-A3B model, they cut the memory footprint from about fifty-seven gigabytes down to about sixteen gigabytes. That means you could run this model on a single high-end consumer GPU instead of a server rack.

Lu: And I think that’s the most impactful part of the paper. They’re not just making a small tweak; they’re providing a complete solution. The algorithm is smart, but the hardware implementation is what makes it practical. They’ve thought about the whole pipeline, from the math to the silicon.

Meng: Exactly. And they even address the practical issues, like shared memory bank conflicts on the GPU. They simulate a modified architecture with more memory banks to make the lookup tables more efficient. That’s the kind of attention to detail that separates a paper that gets cited from a paper that gets deployed.

Tom: And the accuracy holds up, too. We saw that on the reasoning tasks, like GSM8K, they’re getting eighty-six point seven percent accuracy on Qwen3-30B-A3B, compared to fifty point eight percent for QuaRot. That’s a massive jump. So you’re not sacrificing intelligence for speed.

Jane: So we have a method that’s faster, smaller, and more accurate than the alternatives. That’s a rare combination. Let’s bring in Lalam to give us the big-picture cultural impact.

Lalam: Thank you, Jane. This is a significant step towards democratizing AI. When you can run a thirty-billion-parameter model on a single GPU, you open the door for smaller companies, researchers, and even hobbyists to build on top of these powerful models. It means AI assistants, code generators, and creative tools can run locally, which has huge implications for privacy and accessibility. You don’t have to send your data to a cloud server; you can keep it on your device.

Tom: That’s a fantastic point. And it’s not just about cost; it’s about control. We’re moving towards a future where AI is a tool you own, not just a service you rent.

Conclusion: Tom: And that brings us to the end of our discussion on "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts." Jane, let’s wrap this up.

Jane: Gladly, Tom. We started by talking about the problem of outliers in Mixture-of-Experts models, which are a huge bottleneck for compression. Then we walked through CodeQuant’s three-part solution: smoothing the outliers with a learned rotation, clustering the weights with fine-tuned centroids, and reordering the weights to make clustering more effective.

Tom: And we saw the results. Up to four point one five times speedup, massive memory savings, and accuracy that stays close to the original model, even at four-bit precision. This isn’t just a small step; it’s a leap forward for practical AI deployment.

Lu: I’d add that the paper’s strength is in its completeness. It’s not just a theory; it’s a working system that they’ve validated across multiple models and hardware platforms. That’s what makes it so convincing.

Meng: And from an engineering standpoint, the custom kernel design shows that they understand the real-world challenges. It’s one thing to have a good algorithm; it’s another to make it run efficiently on the hardware people actually have.

Lalam: And the cultural impact is profound. By making these models more accessible, we’re enabling a new wave of innovation. Local, private, and efficient AI could become the standard, changing how we interact with technology in our daily lives.

Tom: Well said, everyone. We’re saying goodbye to CodeQuant, but we’re excited to see where this line of research goes. Thanks for listening, and we’ll see you on the next episode.

More episodes

← Home