CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts".
Jane: The paper was written by Xiangyang Yin, Xingyu Liu, Tianhua Xia, Bo Bao, Vithursan Thangarasa et al. from New York University and Cerebras Systems Inc..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone! Today we’re diving into a fresh arXiv paper that’s got me genuinely pumped. It’s called "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts." Jane, that title is a mouthful, but the problem it solves is something we’ve talked about a lot on this channel.
Jane: Absolutely, Tom. And let me just say, for our listeners who aren’t deep in the weeds of machine learning, this paper is about making these massive AI models smaller and faster without breaking their brains. We’re talking about Mixture-of-Experts models, which are like a team of specialists. Instead of one giant brain doing everything, you have a router that sends each question to the right expert. That’s how you get models like Mixtral or Qwen that are huge but still fast.
Tom: Right, and the catch is, when you try to compress these models to run on your phone or a regular server, you run into these things called "outliers." Jane, you want to explain what those are for the folks at home?
Jane: Sure. Imagine you’re trying to take a photo of a room, but one lightbulb is blindingly bright. If you lower the exposure to handle that one bulb, the whole room goes dark. Outliers in AI are like that blinding lightbulb—a few values in the data that are way bigger than everything else. When you try to compress the model, those outliers mess up the whole process.
Tom: And that’s where CodeQuant comes in. The team, from NYU and Cerebras Systems, has a clever way of smoothing out those outliers and then using a technique called clustering to pack the weights more efficiently. We’re going to get into the nitty-gritty in a second, but the headline is that they’re getting up to four point one five times speedup on some hardware while keeping accuracy that blows the other methods out of the water.
Jane: It’s a big deal because it’s not just about making things smaller; it’s about making them usable. We’re talking about running these powerful models on edge devices, which could change how we interact with technology daily.
Tom: Exactly. And I want to bring in our resident expert, Lu, to give us the high-level view. Lu, what’s the big picture here?
Lu: Thanks, Tom. The big picture is that we’re hitting a wall with how much compute and memory these models need. CodeQuant is a smart workaround. Instead of fighting the outliers, they’re essentially redistributing the light in that room so you can take a better photo. They rotate the data to make the outliers less extreme, and then they cluster the weights into groups that are easier to compress. It’s a two-pronged attack that’s proving to be much more effective than the one-size-fits-all quantization we’ve seen before.
Tom: And we’re just getting started. Stick around because we’re going to break down exactly how they do this magic and what it means for the future of AI.
Summary: Jane: Welcome back. We’re still on "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts." Tom, we gave the elevator pitch, but let’s get into the actual summary of what this paper does.
Tom: Right. So, the core problem is that when you quantize a model—that’s the process of using fewer bits to represent numbers—you lose precision. And with outliers, you lose a ton of precision. The authors, led by Xiangyang Yin and Xingyu Liu, propose a three-stage framework to fix this. First, they use something called "Activation-Oriented Outlier Smoothing." That’s their fancy way of saying they apply a learned rotation to the data to make the outliers less extreme before quantization.
Jane: And that’s the "smoothing" part. But they don’t stop there. They also introduce "Adaptive Weight Clustering." Instead of just rounding numbers to the nearest value, they group similar weights together and find the best representative value for each group. It’s like organizing a messy closet by grouping all your t-shirts together and then picking one perfect t-shirt to represent them all.
Lu: And that’s a great analogy, Jane. But the clever part is that they don’t just pick the groups randomly. They fine-tune the centroids—the representative values—to minimize the actual error in the output. They’re not just looking at the weights in isolation; they’re looking at how the weights interact with the activations to produce the final result. That’s what makes their clustering so much better than the standard K-means approach.
Tom: And there’s a third piece called "Permutation-Invariant Outlier Grouping." This one is a bit more technical, but essentially, they reorder the weights before clustering. This makes sure that the extreme outliers are spread out across different groups instead of all ending up in one group and ruining that group’s accuracy.
Jane: So they smooth, they cluster, and they reorder. And the results are pretty stunning. On models like Mixtral eight times 7B, they’re getting perplexity scores—that’s a measure of how confused the model is—down to four point six five, which is incredibly close to the original model’s four point zero one. Other methods were blowing up to sixteen or even seventy-seven.
Lu: That’s the key takeaway. The accuracy is almost preserved, even at four-bit precision. And they show this across a range of models, from smaller ones like Phi-mini-MoE to the massive Qwen3-30B-A3B. The consistency is what’s impressive.
Tom: And we haven’t even talked about the speed yet. That’s coming up in the next segment, so stay tuned.
Improvements: Jane: Welcome back to the show. We’re deep into "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts." Tom, we’ve covered the problem and the method. Now let’s talk about the actual improvements this paper brings to the table.
Tom: Absolutely, Jane. And I want to bring in Meng, our engineer, because this is where the rubber meets the road. The paper doesn’t just stop at the algorithm; they actually built a custom kernel to make this thing run fast on real hardware.
Meng: Yeah, and that’s the part that gets me excited. A lot of papers propose clever algorithms, but they’re too slow to actually use. CodeQuant is different. They designed a lookup-table-based kernel. Instead of doing a bunch of multiplications, they precompute the results and just look them up. It’s like having a multiplication table instead of calculating seven times eight every time you need it.
Tom: And that’s where the speedup comes from. They’re reporting up to two point six three times speedup on an A100 GPU compared to the standard BF16 model. And on a CPU, they’re seeing up to four point one five times speedup. That’s a huge deal for deployment.
Jane: And it’s not just about speed. They’re also saving a ton of memory. On the Qwen3-30B-A3B model, they cut the memory footprint from about fifty-seven gigabytes down to about sixteen gigabytes. That means you could run this model on a single high-end consumer GPU instead of a server rack.
Lu: And I think that’s the most impactful part of the paper. They’re not just making a small tweak; they’re providing a complete solution. The algorithm is smart, but the hardware implementation is what makes it practical. They’ve thought about the whole pipeline, from the math to the silicon.
Meng: Exactly. And they even address the practical issues, like shared memory bank conflicts on the GPU. They simulate a modified architecture with more memory banks to make the lookup tables more efficient. That’s the kind of attention to detail that separates a paper that gets cited from a paper that gets deployed.
Tom: And the accuracy holds up, too. We saw that on the reasoning tasks, like GSM8K, they’re getting eighty-six point seven percent accuracy on Qwen3-30B-A3B, compared to fifty point eight percent for QuaRot. That’s a massive jump. So you’re not sacrificing intelligence for speed.
Jane: So we have a method that’s faster, smaller, and more accurate than the alternatives. That’s a rare combination. Let’s bring in Lalam to give us the big-picture cultural impact.
Lalam: Thank you, Jane. This is a significant step towards democratizing AI. When you can run a thirty-billion-parameter model on a single GPU, you open the door for smaller companies, researchers, and even hobbyists to build on top of these powerful models. It means AI assistants, code generators, and creative tools can run locally, which has huge implications for privacy and accessibility. You don’t have to send your data to a cloud server; you can keep it on your device.
Tom: That’s a fantastic point. And it’s not just about cost; it’s about control. We’re moving towards a future where AI is a tool you own, not just a service you rent.
Conclusion: Tom: And that brings us to the end of our discussion on "CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts." Jane, let’s wrap this up.
Jane: Gladly, Tom. We started by talking about the problem of outliers in Mixture-of-Experts models, which are a huge bottleneck for compression. Then we walked through CodeQuant’s three-part solution: smoothing the outliers with a learned rotation, clustering the weights with fine-tuned centroids, and reordering the weights to make clustering more effective.
Tom: And we saw the results. Up to four point one five times speedup, massive memory savings, and accuracy that stays close to the original model, even at four-bit precision. This isn’t just a small step; it’s a leap forward for practical AI deployment.
Lu: I’d add that the paper’s strength is in its completeness. It’s not just a theory; it’s a working system that they’ve validated across multiple models and hardware platforms. That’s what makes it so convincing.
Meng: And from an engineering standpoint, the custom kernel design shows that they understand the real-world challenges. It’s one thing to have a good algorithm; it’s another to make it run efficiently on the hardware people actually have.
Lalam: And the cultural impact is profound. By making these models more accessible, we’re enabling a new wave of innovation. Local, private, and efficient AI could become the standard, changing how we interact with technology in our daily lives.
Tom: Well said, everyone. We’re saying goodbye to CodeQuant, but we’re excited to see where this line of research goes. Thanks for listening, and we’ll see you on the next episode.
Xiangyang Yin, Xingyu Liu, Tianhua Xia, Bo Bao, Vithursan Thangarasa, Valavan Manohararajah, Eric Sather, Sai Qian Zhang
New York University · Cerebras Systems Inc.
cs.LG
Submitted: 2026-08-14
Updated: 2026-08-18
Code: https://github.com/SAI-Lab-NYU/CodeQuant
Project page: https://apple.github.io/coremltools/docs-guides/source/opt-palettization-overview.html
License: http://creativecommons.org/publicdomain/zero/1.0/
Importance score: 86/100
Key concepts
- Mixture-of-Experts (MoE) Models
- These models operate like a team of specialists rather than one giant brain. A router directs incoming data to the most appropriate 'expert' module, allowing developers to build massive, powerful AI models that remain efficient and fast for use.
- Quantization
- This is the process of making large AI models smaller by representing numbers using fewer bits. While necessary for deployment on consumer hardware, this process can cause significant loss of precision, especially when dealing with extreme data values called outliers.
- Outliers in AI
- In data, outliers are values that are significantly larger or smaller than the rest of the dataset. When compressing a model using quantization, these few extreme outliers can disrupt the entire process and severely degrade the model's overall accuracy.
Terminology
Summary
Summary
This paper introduces CodeQuant, a unified quantization-and-clustering framework designed to address the challenge of activation outliers in low-precision Mixture-of-Experts (MoE) models. The paper states: "Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increasingly central to large-scale language modeling. Under post-training quantization (PTQ), these outliers induce substantial quantization errors, leading to severe accuracy degradation. The authors note that while
recent rotation-based smoothing techniques alleviate the problem by redistributing outlier magnitudes, residual errors remain and continue to impede reliable low-precision deployment."
The paper's core contribution is described as follows: "we tackle this challenge by introducing CodeQuant, a unified quantization-and-clustering scheme that contains smoothing activation outliers via learnable rotation and absorbing weight outliers into fine-tuned cluster centroids for MoE. This design reduces the influence of extreme values by fitting them within cluster centroids, thereby lowering quantization error while maintaining expressive capacity."
The methodology consists of four stages, as outlined in the paper: "Stage 1 applies learnable rotations to smooth activation outliers; Stage 2 permute weight for optimized distribution; Stage 3 introduces clustering fine-tune mechanism to align with objective; and Stage 4 deploys the quantized model using a specialized LUT kernel." The specific contributions are detailed as follows:
-
Activation-oriented Outlier Smoothing (AOS): The paper states,
We first introduce Activation-oriented Outlier Smoothing (AOS), which suppresses activation outliers through rotation matrix adjustment, effectively relocating them into the weight space.
This is achieved by fine-tuning the rotation matrix R via the Cayley transform, with the objective:arg min R XR - Q(XR) 2. The paper explains,By parameterizing the matrix M, this construction guarantees that the matrix R remains orthogonal while keeping the process fully differentiable.
-
Adaptive Weight Clustering and Centroid Finetuning (ACCF): The paper states, "We then propose Adaptive Weight Clustering with Centroid Finetuning (ACCF) and Permutation Invariant Outlier Grouping (POG), which substantially reduce weight quantization error even in the presence of significant outliers." The ACCF method refines grouping and centroid search to minimize output error, with a loss function that includes a KL divergence term for MoE-specific routing preservation:
L = Y - sum i=1 E Π̃ i X̃R W c squared + λD KL(Π̃, Π)for gate and up projections. -
Permutation-Invariant Outlier Grouping (POG): This method reorders weight columns to make the matrix more cluster-friendly. The paper explains, "In Step 1, the variance is computed across the elements within each sub-group. In Step 2, the sub-groups are permuted as indivisible units, ordered by their variance, so as to redistribute high- and low-variance sub-groups more evenly across the larger groups of size g = 4."
-
LUT Kernel and System Implementation: The paper states,
We develop a LUT kernel to demonstrate improvements in hardware efficiency.
This kernel uses a two-level Mux to select outputs from precomputed lookup tables, reducing redundant multiplications. The paper notes,CodeQuant uses a two-level Mux to select the output... By pairing activation and weight for shared-memory access, shared-memory conflicts are reduced compared with separate activation and weight accesses.
The paper evaluates CodeQuant on several MoE models, including Phi-mini-MoE-Instruct, Qwen3-30B-A3B, DeepSeek-V2-Lite, and Mixtral 8x7B. The results show that CodeQuant consistently accelerates inference, lowers memory footprint, and preserves accuracy.
Specifically, the paper reports, Across Phi-Mini-MoE-Instruct, Qwen3-30B-A3B, DeepSeek-V2-Lite and Mixtral 8x7B, CodeQuant consistently accelerates inference, lowers memory footprint, and preserves accuracy.
In the main results, CodeQuant demonstrates substantial improvements over baselines like QuaRot, for instance, on Qwen3-30B-A3B, it reduces perplexity by 5.73 on WikiText2 and 8.52 on C4, while increasing average accuracy by 11.3% compared to QuaRot.
The paper also highlights that CodeQuant achieves up to 4.15× speedup while delivering significantly higher accuracy than state-of-the-art quantization approaches across diverse MoE models.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities.
- Implement a Two-Stage Outlier Smoothing and Weight Clustering Pipeline.
-
Action: I will integrate the
CodeQuantframework into the model compression pipeline. This involves two core modules: -
Activation-Oriented Outlier Smoothing (AOS): I will add a learnable rotation matrix (parameterized via the Cayley transform) to the model's forward pass. This matrix will be fine-tuned on a calibration dataset to minimize the quantization error of the input activations, effectively redistributing outlier magnitudes from activations to weights.
-
Adaptive Weight Clustering and Centroid Finetuning (ACCF): I will replace the standard uniform weight quantization with a clustering-based approach. This involves grouping weights into a set of centroids and then fine-tuning these centroids using a custom loss function that accounts for the model's output (e.g., using a KL-divergence loss on router logits for MoE models). This ensures the clustering process is optimized for the final task, not just for minimizing weight error.
-
Resulting Capability: The AI system can be compressed to 4-bit precision (A4W4) with significantly less accuracy degradation than standard quantization (RTN, GPTQ) or rotation-based methods (QuaRot, SpinQuant). For example, on the Qwen3-30B-A3B model, this pipeline reduces perplexity by 5.73 points on WikiText2 and increases average zero-shot accuracy by 11.3% compared to QuaRot.
- Add a Permutation-Invariant Outlier Grouping (POG) Pre-processing Step.
-
Action: Before the clustering step, I will add a permutation module. This module will analyze the weight matrix and reorder its columns (in sub-groups) to group high-variance and low-variance weights together. This creates a more
cluster-friendly
weight distribution, which is a crucial precondition for the ACCF step to achieve low error, especially when using block-wise quantization. -
Resulting Capability: The system can now effectively use block-wise weight clustering (e.g., with a group size of 1024), which is more hardware-friendly than embedding-wise clustering. This leads to further accuracy improvements under extreme compression, as demonstrated by the performance gains on Phi-mini-MoE-Instruct and DeepSeek-V2-Lite in the paper's A4W4 Block-wise results.
- Deploy a Lookup Table (LUT)-Based Kernel for Inference.
-
Action: I will implement a custom GEMM kernel that leverages the clustered weights and quantized activations. Instead of performing multiply-accumulate operations, this kernel will precompute a lookup table for each weight group and activation value, then use a two-level MUX to fetch the result. This kernel will be designed to store the LUT in shared memory and be optimized to reduce bank conflicts.
-
Resulting Capability: The system can achieve significant inference speedups. The simulation shows an average 2.63× speedup over a BF16 baseline on an A100 GPU. On a CPU, using a similar LUT-based kernel (T-MAC), the system achieves up to a 4.15× speedup over the BF16 baseline, while also reducing memory footprint by over 70%.
The improved AI system will be a highly efficient, low-precision MoE model that is both fast and accurate. Specifically, it can:
-
Run on a single GPU with 4-bit precision while maintaining near-BF16 accuracy on complex tasks like mathematical reasoning (GSM8K, MATH500) and commonsense QA (ARC, HellaSwag, MMLU).
-
Achieve up to 4.15× lower latency compared to a standard BF16 implementation, making it suitable for real-time, edge, or high-throughput serving environments.
-
Reduce its memory footprint by more than 70%, allowing for larger models to be deployed on the same hardware or enabling longer context windows.
-
Maintain stable expert routing in MoE architectures, preventing the degradation in performance that often occurs when the router's decisions are altered by quantization. This is achieved through the KL-divergence penalty in the ACCF loss function.
-
Be robust to extreme compression, maintaining a significant accuracy advantage over other methods even when the weight budget is reduced to 2-bit (A4W2).
Abstract
Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increasingly central to large-scale language modeling. Under post-training quantization (PTQ), these outliers induce substantial quantization errors, leading to severe accuracy degradation. While recent rotation-based smoothing techniques alleviate the problem by redistributing outlier magnitudes, residual errors remain and continue to impede reliable low-precision deployment. In this work, we tackle this challenge by introducing CodeQuant, a unified quantization-and-clustering scheme that contains smoothing activation outliers via learnable rotation and absorbing weight outliers into fine-tuned cluster centroids for MoE. This design reduces the influence of extreme values by fitting them within cluster centroids, thereby lowering quantization error while maintaining expressive capacity. Coupled with a dedicated kernel design for GPU and CPU, CodeQuant achieves up to 4.15 times speedup while delivering significantly higher accuracy than state-of-the-art quantization approaches across diverse MoE models. Our results highlight CodeQuant as a promising direction for efficient and accurate deployment of MoE-based large language models under low-precision constraints. Our code is available at https://github.com/SAI-Lab-NYU/CodeQuant.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Systematic Outliers in Large Language Models
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- QLoRA: Efficient Finetuning of Quantized LLMs
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- DeepGEMM: Accelerated Ultra Low-Precision Inference on CPU Architectures using Lookup Tables
- Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
- Mixtral of Experts
- SqueezeLLM: Dense-and-Sparse Quantization
- Expert-Token Resonance MoE: Bidirectional Routing with Efficiency Affinity-Driven Active Selection
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform
- Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
- HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
- SpinQuant: LLM quantization with learned rotations
- A Closer Look into Mixture-of-Experts in Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks