QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

arXiv:2609.00224 · cs.LG, cs.AI · Submitted 2026-08-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization".

Jane: The paper was written by Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing et al. from University of Notre Dame.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: We’ve talked about the concept, but now let's look at how QTEA actually executes this strategy, because it's quite clever. The paper summarizes its approach by combining ternary quantization with a smart way to handle the inevitable errors.

Jane: It seems they are taking the standard GPTQ-style column-by-column method and making it more robust by incorporating a new layer of logic at each step.

Lu: That's right, Jane; instead of just letting the error accumulate as you move through columns, QTEA introduces specific mechanisms to ensure that the quantization process is stable and accurate across the entire model.

Meng: I was looking at the results, and it seems they manage to achieve an effective one point seven bits per weight while maintaining a high level of accuracy on models like Qwen3-14B. That’s a massive jump in efficiency compared to full FP16.

Lalam: The fact that they are using this highly optimized approach suggests that the future of LLMs is not just about having more parameters, but having smarter, more efficiently packed parameters too.

Tom: So, when we look at the summary of QTEA's performance on Qwen3-14B, we see a sixteen point seven percent relative gain in average zero-shot accuracy over the strongest baseline method. That’s impressive improvement for such heavy compression.

Jane: It’s not just about the final number, Tom; it' about the process of how they are correcting those errors throughout the entire model structure, making sure that every single weight contributes to a stable and accurate representation.

Lu: The combination of ternary encoding and this specific residual handling is what makes QTEA unique in a sub-two-bit PTQ framework.

Meng: It addresses the core challenge of maintaining fidelity when reducing the bit budget by prioritizing those most impactful entries.

Improvements: Tom: QTEA has several technical innovations, and I think understanding these is where it gets really interesting—how they push past previous limitations in quantization.

Jane: One area that stood out was the concept of column-wise rescale refinement, which sounds like they are making the model more adaptive to local weight distributions during training.

Lu: It’s about acknowledging that a static group-wise scale doesn' suboptimal because error propagation changes the local magnitude distribution as you move through columns, so adapting that factor is crucial for accuracy.

Meng: I appreciate the discussion of Error Decay, too; it addresses the order-dependent imbalance in traditional GPTQ where later columns just accumulate errors because there aren't many more to absorb them.

Lalam: That’s a very nuanced point, Meng, and it’s something that really helps me think about how LLMs process information—it ensures the final output isn't dominated by late-stage compensation.

Tom: And then we have the column-wise rescale refinement, which is essentially allowing each column to have its own small scaling factor v j to adjust for this local drift.

Jane: It’s a very sophisticated way of saying that instead of forcing every single column to follow the same global scale, we allow them a little bit of personal space to match their own characteristics after error propagation.

Lu: That adaptive scaling, coupled with Error Decay, ensures that the entire model block remains stable even when dealing with extremely tight bit budgets.

Meng: The fact that this mechanism is integrated into a unified framework shows that these improvements aren' the result of isolated fixes, but a cohesive design approach is what makes it highly practical for real-world deployment.

Conclusion: Tom: We've seen how QTEA works and what its key improvements are, but let’s look at the overall impact on efficiency and performance. The paper presents some very impressive benchmarks across different model sizes.

Jane: It’s fascinating to see the consistent scaling behavior, where QTEA performs strongly whether we are looking at a small model or a massive one.

Lu: I think what we’re seeing is that the fundamental architecture of QTEA—the combination of ternary base and sparse compensation—generalizes across model architectures like Llama and Qwen3.

Meng: The efficiency gains are also worth highlighting, because with this structure, they achieve a seven point two times generation speedup over the standard FP16 baseline when using optimized kernels.

Lalam: This speed increase translates directly into lower latency for users, which is something that can radically change how people interact with and utilize large language models.

Tom: The trade-off between accuracy and compression seems to be where QTEA shines; it provides a strong balance, achieving high accuracy while keeping the model size incredibly small.

Jane: It's not just about the size reduction, Tom; it' about the fact that QTEA manages to achieve this reduction without sacrificing performance on benchmarks like WikiText2 and C4.

Lu: The ability Q to maintain stability in ultra-low-bit quantization is a major technical achievement that will likely inspire many subsequent research methods.

Meng: It’s a practical pathway toward extreme efficiency, proving that sophisticated design can lead to massive improvements in both speed and resource usage.

Conclusion: Tom: As we wrap up today, I want us to summarize the biggest implications of "QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization." It’s a truly remarkable piece of engineering.

Jane: We can say that QTEA offers a viable, highly accurate way to compress large language models by leveraging ternary values and correcting errors, making high-power AI accessible.

Lu: The ability to generalize across diverse model architectures while achieving sub-two-bit precision is a major leap for the theoretical understanding of LLM compressibility.

Meng: The hardware implications are clear; this design enables massive speedups and energy reductions in real deployment scenarios, which is exactly what we need for scalable infrastructure.

Lalam: I hope this technology allows AI to move beyond just research labs and into a more pervasive, efficient presence in everyday life, making information processing faster for everyone.

Tom: That's a great final thought from Lalam. It’s clear that QTEA is not just an academic exercise; it has real-world power.

Jane: We hope to see more of this technology being used to make AI deployment simpler and more efficient across the globe.

Lu: I agree, Jane; we' really need to see this applied further in instruction-tuned or multimodal models as well.

Meng: And I’ll be looking at how these results translate into production-grade systems next time around.

Lalam: Lalam sees a future where the computational cost of creating and running these massive models is significantly reduced, fundamentally changing how we engage with AI.

Tom: Thanks to all of you for sharing your insights on QTEA today, and we’ll be back next time with another paper!

Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi

University of Notre Dame

cs.LG, cs.AI

Submitted: 2026-08-31

Updated: 2026-09-02

Comments: Accepted by EMNLP 2026 Main Conference

Code: https://github.com/Intelligent-Microsystems-Lab/QTEA

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: The paper introduces QTEA, a novel framework for developing highly efficient Large Language Models using ternary representation combined with sparse residual salient weights and by-column

Key concepts

Ternary Quantization
QTEA uses ternary encoding, which is a form of quantization that reduces the precision of model weights. This technique allows for significant compression while maintaining accuracy, contributing to the model's overall efficiency and smaller size.
By-Column Optimization
This method applies techniques similar to GPTQ but refines the process column-by-column. Instead of allowing errors to accumulate across the entire model structure, QTEA incorporates specific logic at each step to ensure stability and accuracy throughout the the quantization process.
Column-wise Rescale Refinement
This is a technique where each column in a layer is allowed its own small scaling factor (v_j). This allows individual columns to adapt to their local weight distributions, preventing errors from propagating and maintaining accuracy despite tight bit budgets.

Terminology

Summary

The paper introduces QTEA, a novel framework for developing highly efficient Large Language Models using ternary representation combined with sparse residual salient weights and by-column optimization. This approach addresses the critical trade-off between model size, computational efficiency, and predictive accuracy, demonstrating substantial improvements over existing quantization techniques while providing a detailed hardware realization of the proposed architecture.

Model Compression and Storage Efficiency

QTEA significantly reduces the total model storage footprint of large models like Llama2-7B. Utilizing a payload structure that consists of approximately 1.7 bpw payload plus 0.3 bpw metadata, QTEA achieves a remarkable compression ratio, reducing the Llama2-7B checkpoint size from 13.48 GB in FP16 down to 2.12 GB, corresponding to a 6.35× storage reduction. This performance places QTEA favorably against other low-bit quantization methods; for instance, compared with 2-bit methods such as GPTQ and SlimLLM, QTEA achieves a smaller model size while maintaining superior accuracy. Furthermore, the ternary representation inherently reduces weight storage: for each 128-column group, five ternary weights are packed into one 8-bit address, resulting in a packed ternary weight storage that is about 9.85× smaller than dense FP16 weight storage.

Ternary-LUT Hardware Acceleration

To realize the computational benefits of the ternary representation, the authors propose a dedicated hardware engine for matrix-vector multiplication (y = xW). This ternary-LUT computation engine is evaluated against a dense FP16 matrix computation baseline. The implementation details highlight several efficiency gains:

  • Compute Logic: The ternary compute logic achieves an area of 28,454.25µ2, which means that before considering SRAM accesses, the ternary compute logic uses about 80% of the dense baseline area (35,573.98µ2).

  • Memory Traffic: The engine replaces repeated dense FP16 weight reads with compact access patterns by reading 26 packed 8-bit weight addresses, thereby replacing dense FP16 weight reads with compact packed-weight reads and LUT accesses.

  • Energy and Latency: End-to-end hardware results show that the ternary engine achieves an average latency speedup of 3.83× over the dense FP16 baseline and an average energy reduction of 69.4%.

Accuracy Improvement Rationale

The superior accuracy achieved by QTEA is rigorously tested through a matched-budget comparison against PT2-LLM. The study aims to determine if the accuracy gain is merely due to allocating more high-precision storage to salient weights. When an iso-memory footprint variant of PT2-LLM is constructed by augmenting it with FP8 column-wise salient weights under the same additional memory footprint, QTEA still maintains a substantial lead. Specifically, QTEA remains 6.08 and 4.87 points higher, respectively, than the augmented baseline on Qwen3-8B and Qwen3-14B. These results suggest that QTEA’s accuracy gain cannot be attributed solely to the additional high-precision salient weights, indicating that the proposed quantization and residual-correction design provides benefits beyond simply increasing the effective precision budget.

Improvements for AI systems

Improvement: Integrate a specialized, hardware-accelerated Ternary Look-Up Table (LUT) Computation Engine into the inference pipeline for large language models (LLMs). This engine must be designed to replace standard dense floating-point matrix multiplication units (FP16 MAC lanes) when processing weights quantized using compact ternary representations.

Specific Functionality:

  • Compute Flow: The system will process matrix-vector multiplications (y = xW) by first calculating scaled activations (xi times vi) via a dedicated FP16 pre-compute MAC, followed by 26 parallel ternary MAC units.

  • Weight Handling: Instead of reading dense FP16 weights, the system reads highly compact packed 8-bit weight addresses. These addresses map directly to the LUTs containing all possible ternary combinations for a given weight block.

  • Performance Gain: This architectural change yields an estimated average **latency speedup of 3.83 times ** and an average **energy reduction of 69.4% ** compared to a dense FP16 baseline, while simultaneously reducing the required compute-logic area by approximately 12% (from 35,573.98 mu squared to 28,454.25 mu squared).

What the Improved AI System Can Do:

The system can execute high-throughput, low-power inference for LLMs (such as Llama 2) on edge devices or power-constrained data centers. It enables running models with significantly reduced MAC operations and dramatically lower memory bandwidth requirements due to the replacement of repeated dense weight reads with compact LUT accesses.


Abstract

Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2 bits. Many leverage unstructured sparsity to mitigate this loss, but at the cost of regularity and GPU-friendly execution. We present QTEA, a sub-2-bit PTQ framework that quantizes weights into ternary values and uses salient weights as residual error compensators. To maintain hardware efficiency, residuals are assigned to selected columns with semi-structured 1:4 sparsity within the salient columns. We further add column-wise rescale refinement to GPTQ-style column-by-column quantization, alternately updating per-column scales and ternary assignments to reduce reconstruction error. We also identify order-dependent error propagation in GPTQ and introduce error decay to attenuate late-stage error accumulation. On Qwen3-14B, QTEA compresses all weights to an effective 1.7 bits per weight while improving average accuracy over the strongest ternary PTQ baseline by 16.7%. It also achieves 1.40 times and 2.61 times lower perplexity on WikiText and C4 respectively. This trend holds on Llama3-8B, where QTEA obtains a 6.6% accuracy gain and 1.34 times / 1.95 times lower perplexity on the same datasets. Finally, we develop a lookup-table based kernel that achieves 7.2 times faster per-token generation over an FP16 baseline. Code is available at https://github.com/Intelligent-Microsystems-Lab/QTEA.

Sources

Related papers