ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization
He-Yen Hsieh, H. T. Kung
Harvard University
cs.LG, cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 16 pages, 8 figures
Code: https://github.com/louisYen/ReRound
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the
Terminology
Summary
ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the LLM. These reconstructed weights act as a guidance signal to disambiguate the rounding direction of weights located close to interval midpoints. To integrate this reconstruction-guided rounding with conventional RTN, ReRound introduces a tolerance metric measuring how far the quantized weight (not the final quantized integer) is away from the midpoint: quantized weights within a tolerance region around midpoints are quantized using diffusion-based reconstructions, whereas weights closer to quantization boundaries are quantized with RTN. By sweeping the tolerance parameter, ReRound generates multiple candidate quantized integer weight matrices and selects the de-quantized weight matrix candidate whose leading singular values most closely match those of the original full-precision weights. This selected candidate determines the tolerance parameter ReRound uses. ReRound is particularly effective for smaller LLMs. Across a range of such models, it consistently outperforms standard RTN for 3-bit and 4-bit weight quantization. ReRound achieves superior accuracy compared to an extensive set of calibration-free methods, remains competitive with calibration-dependent approaches, and operates entirely offline, introducing no additional overhead during low-bit inference. The ReRound strategy represents a new approach for low-bit quantization. The method applies to AI models beyond LLMs. This paper focuses on its applications to small LLMs. Code is available at https://github.com/louisYen/ReRound.
The paper introduces ReRound as a calibration-free framework that revisits midpoint-ambiguous RTN decisions using information learned from the pretrained LLM’s own weights. Quantization maps a range of full-precision values to the same low-bit level, so the quantized integer alone does not reveal the original full-precision value. ReRound trains a conditional diffusion model on paired full-precision and low-bit weight regions from the pretrained model to produce continuous reconstructed weights from each low-bit observation. The reconstructed weights are not deployed as model parameters; instead, they provide additional evidence for whether each original weight should be rounded down or up. The usefulness of this evidence depends on the margin of the original RTN decision. Near the midpoint, where the two assignments have similar scalar error, the reconstructed weight can provide useful guidance. Toward either adjacent quantized integer, RTN has a clearer preference and should be preserved more often. ReRound therefore uses a position-dependent tolerance metric determined by the distance of each quantized weight from the midpoint. This metric permits more reconstruction-guided changes within the midpoint region and increasingly preserves RTN toward the adjacent quantized integers. A change is accepted only when the reconstructed weight favors the assignment opposite to RTN and satisfies this tolerance metric. Varying the tolerance parameter produces a small set of candidate quantized integer weight matrices with different reconstruction-guided assignments. These candidates differ in their rounding patterns across the complete weight matrix. ReRound compares them as complete matrices, and the selected candidate is the one whose de-quantized weight matrix has leading singular values that most closely match those of the original full-precision matrix. This weight-only spectral criterion favors candidates that preserve the dominant structure of the original layer. Both diffusion-based reconstruction and spectral selection require no activation samples, calibration text, or downstream labels. The final model differs from RTN only in selected floor-or-ceiling assignments. The scales, zero-points, group size, and quantized layers remain unchanged. The same procedure can therefore refine standard RTN or be applied to quantization parameters produced by another PTQ method. All diffusion-based reconstruction and spectral selection are performed offline. Because ReRound changes only the quantized integer assignments, it leaves the low-bit representation and inference procedure unchanged.
The contributions of the paper are: introducing ReRound, which trains conditional diffusion on the pretrained model’s weight regions and uses the reconstructed continuous weights to guide midpoint-ambiguous floor-or-ceiling decisions without activation or text calibration data; proposing a position-dependent tolerance metric that controls where reconstruction-guided changes are accepted, together with a weight-only spectral criterion that selects the final quantized integer weight matrix; and demonstrating that ReRound consistently improves RTN at 3 and 4 bits, preserves the same low-bit representation and inference procedure, and can be applied to quantization parameters produced by another PTQ method.
In the method section, ReRound revisits the final rounding assignments under a fixed quantization setup. For each pretrained LLM, a conditional diffusion model is trained once to produce continuous reconstructed weights. These reconstructed weights guide rounding decisions affected by midpoint ambiguity and produce candidate quantized integer weight matrices using a position-dependent tolerance metric. For each matrix, the candidate whose de-quantized weight matrix has leading singular values that best match those of the full-precision matrix is selected. The paper formalizes the midpoint-ambiguous RTN decision: for a weight w in a quantization group with scale Δ and zero-point z, the quantized weight is w̃ = w/Δ + z = l + r, where l = ⌊w̃⌋, u = l + 1, and r ∈ [0, 1). RTN selects the closer adjacent quantized integer. The lower and upper assignments incur errors el(r) = Δ2r2 and eu(r) = Δ2(1−r)2, so flipping the RTN assignment incurs an additional cost cflip(r) = Δ21−2r, which is smallest at the midpoint r = 0.5 and increases toward either adjacent quantized integer. ReRound accordingly gives reconstructed weights greater influence near the midpoint and increasingly preserves RTN away from it.
For conditional weight reconstruction, ReRound trains one conditional diffusion model on local patches from the pretrained LLM’s full-precision weight matrices. The model jointly reconstructs each patch from its low-bit observation, providing a prior learned from the pretrained LLM’s own weights. During training, the low-bit condition is formed by group-wise stochastic quantization: Plow = Qsr(P; Δc, zc), where Δc and zc are the corresponding scale and zero-point. For an entry p ∈ P, the quantized weight is p̃ = p/Δc + zc = l + r, and Qsr selects the quantized integer l with probability 1−r and l+1 with probability r. Only the condition is quantized; the reconstruction target remains the full-precision patch P. ReRound models the conditional distribution pθ(P Plow) using conditional diffusion. At diffusion step t, Gaussian noise ϵ ∼ N(0, I) produces Pt = √ᾱt P + √(1−ᾱt)ϵ, where ᾱt determines the noise level. The denoiser receives Pt, t, and Plow and predicts the added noise, minimizing Ldiff = E[∥ϵ − ϵθ(Pt, t, Plow)∥2]. During diffusion inference, each patch is conditioned on its deterministic RTN quantization: Pcond,k = Qrtn(Pk; Δc, zc). Starting from PT ∼ N(0, I), the conditional reverse process produces Prec,k = Rθ(Pcond,k), where Rθ denotes reverse diffusion from step T to step 0. The reconstructed patches are returned to their original locations to form the reconstructed matrix Wrec. Because each patch is reconstructed jointly, each reconstructed entry depends on its low-bit observation and the surrounding entries in the patch. The reconstructed matrix is used only to guide rounding and is not deployed as the final weight matrix.
For reconstruction-guided rounding, ReRound uses the reconstructed weights to propose alternative floor-or-ceiling assignments under the fixed PTQ quantization parameters. Let w and wrec be corresponding entries of W and Wrec. Using the fixed scale Δ and zero-point z, both are mapped to the same quantized weight space: w̃ = w/Δ + z and w̃rec = wrec/Δ + z. Using l, u, and r, define qrtn = round(w̃), qalt = l + u − qrtn, and qrec = round(w̃rec). The reconstructed weight proposes changing RTN only when qrec = qalt. Whether the proposal is accepted depends on the original quantized weight’s position within its quantization interval. Define its distance to the nearest quantized integer as d(r) = min r, 1−r ∈ [0, 0.5]. Here, d(r) = 0.5 at the midpoint and decreases toward zero near either adjacent quantized integer. The position-dependent tolerance metric for a reconstruction-guided proposal depends on d(r). Given a tolerance parameter τ, ReRound defines τpos(r; τ) = τ exp(−β[(0.5 − δ − d(r))+/(0.5 − δ)]2), where [x]+ = max x, 0. The parameter δ defines the midpoint region r−0.5 ≤ δ, where τpos(r; τ) = τ. Outside this region, β controls how quickly the tolerance metric decreases toward either adjacent quantized integer. The reconstruction deviation, measured in quantization-step units, is ρ(w, wrec) = w̃rec − w̃ = wrec − w/Δ. The candidate assignment for tolerance parameter τ is qτ = qalt if qrec = qalt and ρ(w, wrec) ≤ τpos(r; τ), and qτ = qrtn otherwise. Applying this element-wise produces one candidate quantized integer weight matrix Qτ for each tolerance parameter τ ∈ T.
For candidate selection via spectral preservation, ReRound selects the candidate whose de-quantized weight matrix has leading singular values that most closely match those of the full-precision matrix W. For each tolerance parameter τ ∈ T, the candidate quantized integer weight matrix Qτ is de-quantized using the fixed PTQ parameters: Wτ = Dequant(Qτ; Δ, z). Let m = min dout, din be the smaller dimension of W. ReRound compares the leading k singular values, where k = min m, 128, max(32, ⌊m/16⌋). The spectral discrepancy is Dspec(τ) = ∥σ1:k(Wτ) − σ1:k(W)∥2 / (∥σ1:k(W)∥2 + ϵ), where ϵ > 0 ensures numerical stability. ReRound selects τ⋆ = arg min Dspec(τ) over τ ∈ T. The selected Qτ⋆ provides the final quantized integer assignments and, together with the fixed scales and zero-points, defines the quantized representation of W.
In the experiments, the full baseline comparison includes Gemma 2 2B, Gemma 3 1B, Qwen3 1.7B, OLMo 2 1B, and SmolLM2 1.7B. ReRound is further evaluated on Llama 3.2 1B, Pythia 1.4B, and Phi-2 2.7B. Zero-shot accuracy is evaluated on WinoGrande, PIQA, BoolQ, and SIQA using the LM Evaluation Harness. For SINQ, its quantization scales are adopted and perplexity on WikiText-2 and C4 is also reported. In the main RTN and ReRound experiments, asymmetric uniform weight quantization is used, with W4A16 and W3A16 denoting 4-bit and 3-bit weight quantization with 16-bit activations. All group-wise methods use group size G = 128. For all evaluated quantization methods, the transformer linear layers and the output logit layer (lm head) are quantized. ReRound uses the same scales, zero-points, group size, and layer coverage as group-wise RTN; only selected quantized integer assignments differ. The calibration-free baselines are channel-wise and group-wise RTN, HQQ, BNB FP4, Hadamard-transformed RTN, and CafeQ. BNB FP4 is evaluated only at 4 bits. The 4-bit calibration-based baselines are GPTQ, AdaRound, and SignRound, with GPTQ and AdaRound using 128 calibration samples from C4. For each pretrained LLM, one conditional diffusion model is trained on two GPUs using full-precision patches paired with 2-bit conditions generated by stochastic rounding. Diffusion inference uses one GPU to reconstruct Wrec from deterministic 2-bit RTN conditions. The same reconstructed matrix Wrec is reused for the 3- and 4-bit ReRound runs. For SINQ, separate models are trained for 3-bit and 4-bit conditions. ReRound PTQ uses one GPU to map W and Wrec to the same quantized weight space, construct candidate quantized integer weight matrices over τ ∈ T, and select one candidate for each weight matrix by matching the leading singular values of its de-quantized weight matrix to those of W. It changes at most 1% of the RTN quantized integer assignments in each matrix and uses the same formulation across models.
Comparing with calibration-free PTQ, ReRound improves its matched group-wise RTN baseline for every model at both bit widths, with gains of 0.1–0.9 points at 3 bits and 0.2–1.3 points at 4 bits. At 3 bits, ReRound achieves the highest four-task average on every model. At 4 bits, it achieves or matches the best average on four of the five models. On OLMo 2 1B, HQQ and BNB FP4 obtain averages of 60.4 and 61.1, respectively, compared with 59.4 for ReRound, but ReRound raises its matched group-wise RTN baseline from 58.1 to 59.4, the largest 4-bit gain in the comparison. Comparing with calibration-based PTQ under W4A16 quantization, ReRound achieves the highest four-task average among the quantized methods on both Gemma 2 2B and Gemma 3 1B. The strongest calibration-based method, SignRound, reaches averages of 67.0 on Gemma 2 2B and 59.1 on Gemma 3 1B, whereas ReRound achieves 67.4 and 59.3, respectively, surpassing the best calibration-based result by 0.4 and 0.2 points without using activation or text calibration data.
The ablation on OLMo 2 1B isolates reconstruction guidance, the position-dependent tolerance metric, and spectral selection. Full ReRound achieves the highest four-task average at both bit widths, improving group-wise RTN by 1.3 points at W4 and 0.8 points at W3. Replacing any component lowers the average accuracy. Using SINQ quantization parameters on Qwen3 1.7B, at 4 bits ReRound increases the four-task average from 62.1 to 62.6 and reduces average perplexity from 18.53 to 18.46. At 3 bits, it increases the average from 57.1 to 57.3 and reduces average perplexity from 24.26 to 24.06. For additional models, ReRound improves group-wise RTN for every model at both 3 and 4 bits, with gains ranging from 0.2 to 1.3 points at 4 bits and from 0.1 to 1.6 points at 3 bits. The offline runtime of ReRound is reported: diffusion training is performed once per model on two GPUs, one-GPU diffusion inference produces the reconstructed matrix Wrec, and ReRound PTQ averages 42 to 124 seconds across W3 and W4. All additional computation occurs before deployment; the resulting quantized model uses the same low-bit representation and inference procedure as RTN.
The discussion and limitations note that ReRound revises RTN assignments while keeping scales and zero-points fixed, so it can refine quantization parameters produced by other PTQ methods. However, its candidate set remains constrained by these parameters, motivating joint optimization of quantization parameters and reconstruction-guided rounding. Although the approach may extend beyond LLMs, only small LLMs are evaluated. ReRound also requires a separate diffusion model for each pretrained LLM, introducing model-specific offline cost. Reusing reconstructed weights across bit widths reduces this cost, but scaling to many models remains expensive. In addition, spectral selection operates only in weight space and may not identify the best candidate for every downstream task, while patch-based reconstruction may miss long-range and cross-layer dependencies. Reducing reconstruction cost, transferring weight priors across models, and using broader structural criteria are promising directions. The conclusion states that ReRound is a calibration-free framework that revisits midpoint-ambiguous RTN decisions after quantization parameters are fixed, learns a diffusion prior from the model’s own weights, uses it to guide floor-or-ceiling changes, and selects candidate matrices through spectral preservation. Across diverse LLMs, it consistently improves 3- and 4-bit weight-only PTQ without activation or text calibration data. Because ReRound changes only selected quantized integer assignments, the quantization parameters, low-bit representation, and inference procedure remain unchanged. These results show that structural patterns in pretrained weights can serve as an effective prior for resolving ambiguous rounding decisions.
Improvements for AI systems
Improvements to AI Systems Based on ReRound:
-
Calibration-Free Low-Bit Quantization for Edge Deployment: AI systems can now achieve 3-bit and 4-bit weight quantization without requiring activation samples, calibration text, or downstream labels. This enables deployment of smaller LLMs on resource-constrained devices (e.g., mobile, IoT) with minimal accuracy loss (0.1–1.6 points improvement over RTN across models), while maintaining identical inference speed and memory footprint as standard RTN.
-
Diffusion-Guided Weight Reconstruction as a General Prior: The system learns a conditional diffusion model from the pretrained model’s own weight distributions, capturing structural patterns (e.g., local correlations, spectral properties) that serve as a reusable prior. This prior can be applied to any PTQ method’s quantization parameters, not just RTN, allowing for iterative refinement of existing quantization pipelines without retraining the base model.
-
Position-Dependent Tolerance for Adaptive Rounding Decisions: The system introduces a tolerance metric that dynamically balances reconstruction-guided changes near quantization midpoints with RTN preservation near boundaries. This reduces rounding errors by up to 1.3 points in accuracy (e.g., OLMo 2 1B at 4-bit) while changing fewer than 1% of quantized integer assignments, ensuring minimal disruption to the original weight structure.
-
Spectral-Preservation-Based Candidate Selection: By comparing leading singular values of de-quantized candidate matrices against the full-precision original, the system selects the rounding pattern that best preserves the layer’s dominant structural information. This weight-only criterion avoids costly activation-based evaluation and works offline, making it suitable for automated model compression pipelines.
-
Cross-Bit-Width Reuse of Reconstructed Weights: The system trains a single diffusion model per LLM (using 2-bit stochastic rounding conditions) and reuses the reconstructed weight matrix for both 3-bit and 4-bit quantization. This reduces offline training cost by 50% and enables rapid experimentation across bit widths without retraining.
-
Post-Hoc Refinement of Any PTQ Method: The system can take quantization parameters (scales, zero-points, group sizes) from any existing PTQ method (e.g., SINQ, GPTQ) and improve their accuracy by revising only the integer assignments. This makes it a drop-in enhancement layer for existing quantization workflows, improving perplexity by up to 0.2 points (e.g., Qwen3 1.7B with SINQ at 3-bit) without changing the deployment format.
-
Scalable Offline Optimization for Model Compression: The entire ReRound process (diffusion training + inference + PTQ refinement) runs offline, with PTQ refinement taking only 42–124 seconds per model on a single GPU. This enables automated model compression pipelines to iterate over multiple quantization configurations efficiently, without needing to re-run expensive calibration or fine-tuning steps.
What the Improved AI System Can Do:
-
Deploy small LLMs (1–3B parameters) at 3-bit or 4-bit weight precision on edge devices with accuracy comparable to or better than calibration-based methods (e.g., SignRound, GPTQ) while using zero calibration data.
-
Automatically refine any existing quantized model (from RTN, HQQ, GPTQ, SINQ, etc.) by adjusting rounding decisions to better preserve weight structure, improving zero-shot reasoning (WinoGrande, PIQA, BoolQ, SIQA) and language modeling perplexity.
-
Generate multiple candidate quantized models and select the one with the best spectral fidelity to the original, enabling trade-off analysis between bit-width, accuracy, and deployment constraints.
-
Operate fully offline, making it suitable for privacy-sensitive applications where activation data cannot be collected or shared.
Abstract
ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the LLM. These reconstructed weights act as a guidance signal to disambiguate the rounding direction of weights located close to interval midpoints. To integrate this reconstruction-guided rounding with conventional RTN, ReRound introduces a tolerance metric measuring how far the quantized weight (not the final quantized integer) is away from the midpoint: quantized weights within a tolerance region around midpoints are quantized using diffusion-based reconstructions, whereas weights closer to quantization boundaries are quantized with RTN. By sweeping the tolerance parameter, ReRound generates multiple candidate quantized integer weight matrices and selects the de-quantized weight matrix candidate whose leading singular values most closely match those of the original full-precision weights. This selected candidate determines the tolerance parameter ReRound uses. ReRound is particularly effective for smaller LLMs. Across a range of such models, it consistently outperforms standard RTN for 3-bit and 4-bit weight quantization. ReRound achieves superior accuracy compared to an extensive set of calibration-free methods, remains competitive with calibration-dependent approaches, and operates entirely offline, introducing no additional overhead during low-bit inference. The ReRound strategy represents a new approach for low-bit quantization. The method applies to AI models beyond LLMs. This paper focuses on its applications to small LLMs.
Sources
- SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Gemma 3 Technical Report
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- Quantized Compressed Sensing with Score-based Generative Models
- 2 OLMo 2 Furious
- CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
- Learning Single Index Models with Diffusion Priors
- Neural Network Diffusion
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks