QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction".
Jane: QUASAR introduces a Quantization-Aware Training (QAT) method that continuously performs lightweight, loss-aware reconstruction during training to lower the loss floor and improve model quality.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Moving on to the conclusion of "QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction," we're looking at how this work fits into the broader landscape of model compression. The authors, Vincent Counathe, Ben Athiwaratkun, Christopher De Sa, and Tianyi Zhang from Cornell University and Together AI, have presented a method that continuously performs loss-aware reconstruction throughout QAT to reduce the loss floor <ref:2608.13966#pg0>.
Jane: The authors argue that this continuous process is necessary because performing the reconstruction once on a frozen model isn't feasible when weights are evolving during QAT, as they pointed out in their motivation <ref:2608.13966#pg0>. They show that this approach addresses the structural mismatch between the loss calculation and the weight updates by optimizing for reconstruction error at every step <ref:2608.13966#pg2>.
Lu: What's fascinating is how they use an exponential moving average of squared gradients to approximate the Hessian, which gives them those per-parameter saliency scores <ref:2608.13966#pg0>. That technique is really clever for estimating the importance of each weight relative to the loss landscape <ref:2608.13966#pg2>.
Meng: From an engineer's view, it’s powerful that they manage to keep this process lightweight while achieving improvements across various bit widths and model families <ref:2608.13966#pg2>. I'm curious if the computational cost of calculating those saliency scores scales predictably with the model size, because if it does, we can plan for deployment environments better <ref:2608.13966#pg0>.
Lalam: If we can consistently achieve lower evaluation loss, like reducing it by at least twenty-nine percent at INT2 compared to standard QAT <ref:2608.13966#pg2>, that translates directly into deploying more capable AI systems <ref:2608.13966#pg1>. This kind of reliable quality boost makes our AI much more useful for complex tasks <ref:2608.13966#pg2>.
Tom: So, to wrap up this paper, QUASAR provides a way to directly influence both the training path and the final model loss by focusing on minimizing that reconstruction error <ref:2608.13966#pg0>. It’s not just about a single post-training step anymore; it’s integrated into the training itself <ref:2608.13966#pg0>.
Jane: And the implications are significant because this technique shows that continuous, loss-aware reconstruction can effectively control the final quality of quantized models <ref:2608.13966#pg2>. It moves us closer to training models precisely for the precision they'll actually be serving at <ref:2608.13966#pg1>.
Lu: I see this as a way to make low-precision deployment truly principled rather than just a brute-force approach <ref:2608.13966#pg0>. The theoretical result linking the reconstruction error directly to the convergence bound is quite compelling <ref:2608.13966#pg2>.
Meng: I still wonder about the practical limitations they mention; specifically, they note that for supervised fine-tuning on models like Qwen3-4B Base at INT4/three/two using reasoning traces, QUASAR outperforms competitive methods by at least ten point nine points in average accuracy across five math benchmarks <ref:2608.13966#pg2>. That's a solid metric, but does that performance gain hold up when we move to much larger models?
Lalam: It holds up well for the specific tasks they tested, which is encouraging for our current reasoning applications <ref:2608.13966#pg1>. If this technique scales effectively, it could unlock new levels of intelligence in smaller, more accessible AI tools <ref:2608.13966#pg2>.
Tom: So, the authors have shown that QUASAR consistently yields lower training and evaluation loss than other competitive QAT baselines across different bit widths and model families <ref:2608.13966#pg2>. It really seems like a solid contribution to making low-bit AI practical <ref:2608.13966#pg0>.
Jane: That's the core finding, Tom, that the continuous minimization of loss-aware reconstruction error leads to tangible improvements in model quality and training stability <ref:2608.13966#pg2>. It’s a method that provides a direct lever on both where the training goes and how good the final quantized model ends up being <ref:2608.13966#pg0>.
Lu: I think this work opens up new avenues for research into how to design loss functions that better capture the quantization noise effect <ref:2608.13966#pg1>. We need to explore what other types of online surrogates we can use besides the exponential moving average <ref:2608.13966#pg0>.
Meng: For now, I'll be focused on how we can integrate their saliency estimation into our existing pipeline without introducing major new dependencies or slowing down the iteration cycle significantly <ref:2608.13966#pg0>. We need concrete performance data on that trade-off <ref:2608.13966#pg2>.
Lalam: I'm really optimistic about this; if we can deploy these models more reliably, it will foster a culture where we push the boundaries of what’s possible with quantized AI <ref:2608.13966#pg2>. It gives us a strong foundation for building better tools <ref:2608.13966#pg1>.
Tom: Well, that covers the main points from "QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction." We've seen how it tackles the loss floor problem by using continuous reconstruction error minimization <ref:2608.13966#pg0>. Next up, we’ll look at the deeper implications of this work for deployment and future AI development.
Conclusion: Tom: So we've been diving deep into QUASAR, and now it’s time to wrap up by talking about what this paper is all about and why it matters for the broader world of AI.
Jane: Exactly, Tom; we need to get back to those basics—what is this paper actually trying to achieve in simple terms?
Lu: It’s fundamentally about tackling a persistent problem in quantization-aware training, specifically that annoying loss floor that keeps preventing models from getting as good as they could be.
Meng: From my side, I'm focused on the practical implications; what does this mean for how we actually deploy these quantized models in real-world scenarios?
Lalam: For me, it points toward a future where AI systems are trained to perform at the exact level of precision they will actually use in production.
Tom: That’s the big picture, Lalam; so QUASAR is essentially an iterative training method that keeps correcting itself based on how the quantization affects the loss every single step.
Jane: Right, it’s like having a constant, intelligent feedback loop during training that adjusts to keep the final result as high quality as possible for a given bit width.
Lu: The authors basically found that by minimizing this specific reconstruction error continuously, you establish a much better target for both how the model learns and what its final quantized performance will be.
Meng: So, if we can reliably lower that floor without drastically increasing training time or adding complex new inference steps, that’s a huge win for deployment efficiency.
Lalam: That reliability is key; it means we can trust these smaller models to perform consistently in critical reasoning tasks because the training process itself is optimized for that specific hardware constraint.
Tom: It sounds like QUASAR gives us a direct lever on both the training trajectory and the final model quality, which is really significant for anyone working on model optimization.
Jane: That's right; it moves us away from a one-off post-training fix toward a more integrated approach that ensures better quality from the very beginning of the QAT process.
Lu: I think this work opens up new avenues for how we design loss functions that truly capture the noise introduced by quantization, which is where the real creative potential lies for future AI research.
Meng: I'm still curious about how they managed to keep that continuous reconstruction process lightweight enough not to slow down our standard training pipelines significantly.
Lalam: That efficiency is what makes this method so appealing; it suggests we can achieve higher quality results without needing massive computational resources just for the optimization process itself.
Tom: So, QUASAR is about using a continuous, loss-aware reconstruction error minimization as the principled target to control both training and final quantized loss.
Jane: That’s a perfect summary of the core idea; it’s an elegant way to manage that structural mismatch we discussed earlier during QAT.
Lu: What this implies for the wider AI community is that we can start designing quantization strategies with much more explicit quality targets baked into the training mechanism itself, rather than treating quantization as just another post-processing step.
Tom: Absolutely; it shows a principled way to approach low-bit deployment that prioritizes actual performance gains over just achieving a certain bit width.
Jane: And the authors’ work sets a new benchmark for what continuous optimization can achieve within the constraints of model compression techniques.
Cornell University · Together AI
cs.LG, cs.CL, stat.ML
Submitted: 2026-08-14
Updated: 2026-10-05
Comments: 44 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: QUASAR introduces a Quantization-Aware Training (QAT) method that continuously performs lightweight, loss-aware reconstruction during training to lower the loss floor and improve model quality.
Key concepts
- Loss Floor Gap
- QAT often has a higher loss floor than full-precision training because the forward pass uses quantized weights while the optimizer updates full-precision weights. This mismatch causes poor training trajectories and limits achievable accuracy.
- Loss-Aware Reconstruction Error
- This error measures how poorly the model's reconstructed weights align with its actual performance loss. QUASAR minimizes this error continuously during training to guide the model toward better quantization outcomes.
- Saliency Scores
- These scores estimate how much each individual weight influences the overall loss function. They are derived by approximating the Hessian using an exponential moving average of squared gradients, helping QUASAR decide which weights need adjustment.
- Quantization and Dequantization Stages
- QUASAR splits reconstruction into two parts: quantization, which maps full weights to discrete codes based on a clipping range, and dequantization, which maps those codes back to reconstructed weights using learned scale and offset parameters.
Terminology
Summary
QUASAR introduces a Quantization-Aware Training (QAT) method that continuously performs lightweight, loss-aware reconstruction during training to lower the loss floor and improve model quality. The core finding is that by minimizing the loss-aware reconstruction error, QUASAR establishes a principled optimization target that controls both the training trajectory and the final quantized model's loss.
The Gist
QUASAR introduces a QAT method that continually minimizes loss-aware reconstruction error to improve the training trajectory and lower the loss floor.
Motivation: The Loss Floor Gap in QAT
Quantization-aware training (QAT) suffers from a structural mismatch where the forward pass and loss use quantized–dequantized reconstruction weights, while the optimizer updates latent full-precision weights. This mismatch leads to suboptimal training trajectories and a higher loss floor compared to full-precision training. Prior methods mitigate this by minimizing loss-aware reconstruction error, but doing so once for a frozen model is impractical during QAT because the latent weights evolve at every step. QUASAR addresses this by continuously performing lightweight, loss-aware reconstruction in the training loop.
QUASAR Methodology
QUASAR decomposes reconstruction into two stages: quantization and dequantization. To make the loss-aware objective tractable, it approximates the Hessian using an exponential moving average of squared gradients to yield per-parameter saliency scores. It then searches over a small set of candidate clipping ranges, each defining a different assignment of discrete codes; for each assignment, it solves for the dequantization parameters in closed form via saliency-weighted least squares and selects the candidate with the smallest loss-aware reconstruction error. This process is summarized in Algorithm 1.
Theoretical Analysis: Bounding Convergence
The theoretical analysis decomposes the QAT convergence bound into three terms: initialization, minibatch noise, and loss-aware reconstruction error. The analysis shows that only the last term depends on the reconstruction map and is exactly the objective QUASAR minimizes at each step. Theorem 1 proves that QUASAR’s reconstruction-induced gradient mismatch satisfies a bound controlled by minimizing the minimum loss-aware reconstruction error: QUASAR’s reconstruction-induced gradient mismatch satisfies∥∇L(rt) − ∇L(wt)∥2 ≤ C Sbt(rt) = C min r∈Rt Sbt(r).
Empirical Results and Performance
Empirically, QUASAR consistently achieves lower training and evaluation loss than competitive QAT baselines across bit widths. For Qwen3-4B-Thinking-2507 at INT2, QUASAR reduces KL divergence by about 30% compared to Standard QAT. Furthermore, when applied to supervised fine-tuning for mathematical reasoning, QUASAR outperforms both QAT and full-precision training followed by PTQ by at least 10.9 percentage points across five math benchmarks at INT2. Across NVFP4, QUASAR reduces held-out KL by approximately 30% relative to standard QAT while also improving downstream accuracy. The method increases training step time by only about 1.5% and introduces no inference overhead.
Key Components and Trade-offs
QUASAR’s saliency estimates are derived from the exponential moving average of squared gradients, which tracks the diagonal curvature. The search over clipping ranges determines code assignments, while the weighted least-squares fit determines dequantization parameters. Ablation studies confirm that reconstruction must remain part of the training loop for improvement and that using Adam’s second moment as saliency performs on par with maintaining a separate Fisher estimate at lower computational cost. The resulting low-bit model deploys exactly as an RTN-quantized one, with no inference-time changes or overhead.
Conclusion
QUASAR successfully mitigates the mismatch between reconstructed weights and latent weights by optimizing the loss-aware reconstruction error throughout training, providing a direct lever on both the training trajectory and the final quantized model's loss. It consistently achieves superior quality compared to existing methods across various bit widths and model families.
(Note: The summary adheres strictly to the provided text, focusing only on QUASAR's mechanism, theoretical foundation, and empirical results as requested.)
How it works
QUASAR decomposes reconstruction into two stages: quantization, which maps full-precision weights to discrete codes while treating the clipping range as a free parameter, and dequantization, which maps those codes back to reconstructed weights using a learned scale and, for asymmetric quantization, an optional offset. To make the loss-aware objective tractable during training, QUASAR approximates the Hessian using an exponential moving average of squared gradients, yielding per-parameter saliency scores that estimate each weight’s effect on the loss.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the QUASAR paper thoroughly. The core contribution is providing a principled, continuous optimization target (minimizing loss-aware reconstruction error) for Quantization-Aware Training (QAT), which effectively lowers the loss floor
of low-bit models.
Here are the specific improvements to AI systems that can be achieved by implementing QUASAR:
The implementation of QUASAR enables the creation of highly efficient, high-quality language models across various precision formats, leading to four primary system capabilities:
- Enhanced Model Fidelity in Low-Bit Inference (Lower Loss Floor):
QUASAR directly addresses the structural mismatch between quantized weights and the loss function during QAT by continuously minimizing the loss-aware reconstruction error. This ensures that as latent weights evolve during training, the reconstructed forward pass remains as close as possible to what a full-precision model would produce.
The improved system can achieve:
A. Lower held-out KL divergence to the full-precision teacher (up to 29% reduction at INT2).
B. Higher average accuracy across eight downstream benchmarks (e.g., GSM8K, MMLU) compared to Standard QAT baselines, particularly at lower bit widths (INT2 and INT3).
C. A final quantized model whose loss is directly controlled by the training objective, ensuring the deployed artifact maintains high quality even when compressed aggressively.
- Superior Knowledge Acquisition in Quantized Domains (Effective Adaptation):
QUASAR's ability to minimize reconstruction error during QAT allows models to learn new capabilities directly within their low-bit weight constraints, surpassing methods that require a lossy post-training conversion after full fine-tuning.
The improved system can achieve:
A. Significantly better performance on mathematical reasoning tasks (e.g., GSM8K, MATH benchmarks), where QUASAR outperforms competitive QAT and PTQ baselines by at least 10.9 percentage points at INT2.
B. Sustained reasoning capabilities over long contexts (e.g., HMMT-2025), as the reconstruction error minimization ensures that the adapted behavior is preserved deeper into generated traces, preventing capability degradation during inference/adaptation phases.
- Format Agnostic Deployment (Cross-Platform Compatibility):
QUASAR's methodology is designed to be format-agnostic, supporting not only standard integer quantization (INT2, INT3, INT4) but also production floating-point formats like NVFP4.
The improved system can achieve:
A. Direct deployment on specialized hardware (like Blackwell GPUs) using native 4-bit floating-point formats (NVFP4), retaining the quality gains achieved through QAT distillation.
B. A unified training pipeline that can be seamlessly adapted to any target precision format by simply modifying the reconstruction parameters in the forward pass, without requiring a complete retraining cycle for each new bit width.
- ⏱ Optimized Training Efficiency (Minimal Overhead):
QUASAR maintains a high level of training efficiency by using an online Hessian proxy (Exponential Moving Average of squared gradients) and performing lightweight, closed-form optimization for dequantization parameters instead of costly per-step full-Hessian calculations or expensive second-order PTQ reconstructions.
The improved system can achieve:
A. Training step time increases by only about 1.5% compared to Standard QAT, making the continuous optimization feasible for large models without prohibitive computational costs.
B. No inference-time overhead is introduced, ensuring that the model deploys exactly as a standard low-bit quantized artifact (e.g., RTN or E4M3), maximizing deployment throughput on specialized accelerators like NVIDIA’s Rubin LUT formats.
Sources
- On the Convergence of SGD with Biased Gradients
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Scaling Law for Quantization-Aware Training
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- A Survey of Quantization Methods for Efficient Neural Network Inference
- The Llama 3 Herd of Models
- Scaling Laws for Precision
- The Perfect Blend: Redefining RLHF with Mixture of Judges
- Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
- Let's Verify Step by Step
- LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
- AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
- Pretraining Large Language Models with NVFP4
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Microscaling Data Formats for Deep Learning
- Qwen3 Technical Report
- Instruction-Following Evaluation for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks