Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling".
Jane: Quantization noise can accumulate over diffusion model denoising trajectories, degrading generation quality, and this paper introduces Q-Drift,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's start by looking at the title and who came up with this work. It’s "Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling," and it was written by Sooyoung Ryu, Mathieu Salzmann, and Saqib Javed. The name itself suggests they are focusing on correcting a drift issue specifically related to quantization in diffusion sampling methods.
Jane: That title tells us exactly what the core problem is: we're dealing with a specific type of drift that happens when you quantize models used in diffusion sampling, so it’s very targeted research. It’s about keeping the sampling process aligned with what a full-precision model would do, even under compression.
Lu: The authors are coming from institutions like Seoul National University and EPFL, which suggests they bring a strong theoretical background to this problem of modeling stochastic perturbations in continuous time dynamics, which is where diffusion sampling lives.
Meng: I've seen work on quantization stabilization before, but what makes this paper different is its focus on the trajectory drift rather than just stabilizing the output at a single step; that's a crucial distinction for deployment.
Lalam: I think it’s important to hear about their approach because if they can provide a principled way to adjust the drift, it gives us more confidence in pushing these models into real-world applications where storage and inference speed are critical constraints.
The paper's summary: Tom: So, what is the main gist of what "Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling" actually proposes? Essentially, they introduce Q-Drift as a sampler-side correction that treats the quantization error as an implicit stochastic perturbation at every denoising step to derive an adjustment that keeps the marginal distribution preserved.
Jane: In simpler terms, it means instead of just fixing the image quality at one point in time, they are modeling how that noise evolves over the entire sequence of steps, which is what causes the degradation we see in practice. They derive a specific formula for this drift adjustment based on empirical observations from models like D2-DPM.
Lu: The methodology involves modeling the quantization error as an implicit Gaussian perturbation injected into the denoiser output at each step, then using a generalized SDE framework to reparameterize the marginal-preserving family in terms of a noise level sigma(t).
Meng: That sounds mathematically intensive; I need to know if this is something that can be implemented without requiring an overhaul of the core sampling loop or adding significant computational overhead during inference.
Lalam: The paper summarizes that they estimate a timestep-wise variance statistic from calibration, and they show that this correction factor is plug-and-play with common samplers and PTQ methods, which makes it very appealing for immediate practical use.
The paper's improvements: Tom: Moving on to the specific improvements the authors detail, they focus on deriving a correction factor, c i, based on matching an implicit stochastic increment variance to the diffusion term in the Euler–Maruyama discretization for a first-order sampler.
Jane: They show that this leads to a final Q-Drift update formula where the next state is calculated by scaling the deterministic Euler update component with (one + c i), which explicitly incorporates that derived correction factor into the sampling process.
Lu: The derivation involves matching this implicit variance to the diffusion term, resulting in c i = sigma i / (two sigma i) V sigma i, where V sigma i is estimated during an offline calibration phase. This shows they are deriving a statistically grounded correction rather than just guessing what might work.
Meng: The calibration requirement is interesting; they state that this variance statistic can be estimated with as few as five paired full-precision and quantized calibration runs, which makes the setup incredibly practical for deployment environments where extensive data collection is not feasible.
Lalam: And they verified that these estimated correction factors stay stable even when you use nested subsamples of a standard 5K-prompt calibration run, which speaks to the robustness of their statistical estimation process.
Conclusion: Tom: So, to wrap up what we've discussed about "Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling," the main point is that this method provides a principled way to correct for the drift caused by accumulating quantization noise during diffusion sampling trajectories.
Jane: They show that by treating the quantization error as an implicit perturbation, we can derive a marginal-distribution-preserving adjustment that works across various samplers and architectures without needing to change the underlying model or quantization scheme.
Lu: The paper’s extension of this drift-rescaling principle to other samplers, like Flow-matching and DPM-Solver++, is quite telling; it suggests a general mathematical principle governing how we can handle this specific type of error correction in diffusion dynamics.
Meng: From an engineering perspective, the negligible overhead at inference mentioned by the authors is what makes me most optimistic; if we can get this kind of fidelity improvement without adding significant latency, it becomes much more viable for real-time applications.
Lalam: I think the ultimate implication here is that we can deploy these large diffusion models in production pipelines where the primary challenge was previously treating quantization error as purely local, leading to a consistent visual quality improvement across diverse models.
Department of Computer Science and Engineering, Seoul National University · School of Computer and Communication Sciences, EPFL · Meta Reality Labs
cs.CV, cs.LG
Submitted: 2026-03-18
Updated: 2026-09-28
Comments: 21 pages, 3 figures, 7 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: Quantization noise can accumulate over diffusion model denoising trajectories, degrading generation quality, and this paper introduces Q-Drift, a sampler-side correction that treats quantization
Key concepts
- Quantization Noise Accumulation
- When large diffusion models are deployed using post-training quantization, small errors occur at each step. These errors do not cancel out; instead, they accumulate along the entire denoising trajectory. This accumulation degrades the final image quality because the model deviates from the true data distribution.
- Implicit Stochastic Perturbation
- The paper models quantization error as an invisible noise injected into every denoising step, similar to a stochastic process. Instead of treating it as a fixed error, Q-Drift treats this noise implicitly during sampling to derive a correction that accounts for its cumulative effect on the distribution.
- Marginal-Distribution-Preserving Drift Adjustment
- The goal is not just to fix errors at one step but to ensure the overall distribution of generated samples remains accurate. Q-Drift calculates a specific adjustment factor based on this noise model, which scales the sampler's movement to counteract the drift caused by quantization error.
- Q-Drift Update Formula
- The final correction involves scaling the deterministic update step: x(n+1) = x(n) + Δσ_i * (1 + c_i)⊙ ϵˆ. This formula incorporates a correction factor (c_i) derived from the noise variance, effectively adding a controlled stochastic component to the standard sampling process.
Terminology
Summary
Quantization noise can accumulate over diffusion model denoising trajectories, degrading generation quality, and this paper introduces Q-Drift, a sampler-side correction that treats quantization error as an implicit stochastic perturbation to derive a marginal-distribution-preserving drift adjustment.
The gist
Q-Drift is a principled sampler-side correction that treats quantization error as an implicit stochastic perturbation on each denoising step and derives a marginal-distribution-preserving drift adjustment.
Motivation and Problem Statement
Post-training quantization (PTQ) is used to deploy large diffusion models, but quantization noise can accumulate over the denoising trajectory and degrade generation quality. A key limitation of most PTQ methods is that they treat quantization error as a step-local error independent across sampling steps, whereas diffusion proceeds iteratively, allowing small local errors to accumulate along the sampling trajectory. This suggests that corrections should be designed to preserve marginal distribution rather than only enforcing per-step denoiser fidelity.
Methodology: Modeling Quantization Noise
The authors model quantization error as an implicit Gaussian perturbation by following empirical observations from D2-DPM, which shows that both the quantized output and the quantization noise are well-approximated by Gaussian distributions at each timestep. They define the quantization noise as:
“the quantization noise as implicit noise injected into the denoiser output at each step.”
They adopt a generalized SDE framework where sampling corresponds to reverse-time dynamics, which can be augmented with Langevin diffusion processes controlled by a function of time, β(t). They then reparameterize the marginal-preserving generalized SDE family in terms of the noise level σ(t) to obtain an equation that controls both injected noise and paired deterministic noise decay.
Derivation for Euler Sampler
The derivation focuses on a first-order Euler sampler. By reparameterizing the marginal-preserving SDE family, they show that quantization error introduces an implicit stochastic increment with one-step variance proportional to the step size:
“the last term behaves as an implicit stochastic increment with one-step variance (∆σi)2Vσi.”
They match this implicit variance to the diffusion term in the Euler–Maruyama discretization, deriving a correction factor, denoted as c i:
c i = Δσi / (2σ i) Vσ i.
Final Q-Drift Update
This correction factor is then used to scale the deterministic Euler update component of the sampler. The final Q-Drift update is:
**“x(n+1) = x(n) + Δσ i **
(1 + c i)⊙ ϵˆ”
Calibration and Efficiency
The correction factor Vσi is estimated during an offline calibration phase. The statistic is defined as the expected variance of the quantization noise conditional on the quantized output:
Vσi = E[Var(Δϵ i ˆϵθ(x, σ i, c))]. The paper demonstrates that this calibration remains effective with as few as 5 paired full-precision/quantized calibration runs.
Furthermore, they verify that the estimated correction factors are stable across nested subsamples of a standard 5K-prompt calibration run.
Results and Validation
Empirically, Q-Drift improves FID over quantized baselines across six diverse text-to-image models (spanning DiT and UNet), three samplers (Euler, flow-matching, DPM-Solver++), and two PTQ methods (SVDQuant, MixDQ). The method consistently recovers a meaningful fraction of quantization-induced performance degradation
while preserving CLIP scores. The correction incurs negligible overhead at inference.
Comparison with a D2-DPM style bias correction shows that the latter is often detrimental, whereas Q-Drift improves distributional fidelity. Qualitative visual comparisons show that Q-Drift yields subtle but consistent improvements in fine details or structure relative to the quantized baseline.
The model relies on assumptions such as element-wise independence and isotropic parameterization for covariance blocks, which are deemed calibration-friendly approximations.
Extension to Other Samplers
The drift-rescaling principle is shown to be applicable beyond the Euler sampler. They provide derivations for extending Q-Drift to other samplers, including the Flow-matching sampler (where the correction factor is applied directly to the velocity field) and DPM-Solver(++), by matching their respective SDE forms in either x or y space. This demonstrates that the same drift-rescaling principle can be derived for other samplers.
Conclusion
Q-Drift is presented as a drop-in method agnostic to sampler, architecture, and baseline PTQ method
that generally improves FID over quantized baselines while keeping CLIP scores comparable.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems using Q-Drift, and what these improved systems can achieve:
The core improvement offered by Q-Drift is a mechanism to restore the intended sampling distribution when diffusion models are deployed with Post-Training Quantization (PTQ). This addresses the drift
caused by accumulating quantization noise throughout the iterative denoising process.
Here are the specific improvements and capabilities:
-
The system can maintain high generation quality (measured by FID and CLIP scores) even when running a large diffusion model under aggressive low-bit quantization settings (e.g., 4-bit or lower).
-
It allows for the deployment of highly compressed diffusion models without significant degradation in visual fidelity compared to full-precision models.
-
The correction mechanism is
plug-and-play,
meaning it can be seamlessly combined with existing PTQ methods (like SVDQuant or MixDQ) and common samplers (Euler, DPM-Solver++, Flow-matching) without requiring any modifications to the model architecture or the quantization scheme itself. -
It introduces only negligible overhead during inference, as the correction relies on a small set of precomputed calibration runs and a lightweight, step-wise drift rescaling factor that is applied at each denoising step.
-
The system can be calibrated efficiently using only as few as 5 paired full-precision/quantized calibration runs to estimate the necessary step-wise variance statistics, making it practical for deployment where extensive training or calibration is infeasible.
In summary, the improved AI systems can:
-
Generate high-fidelity images from text prompts using quantized diffusion models (e.g., SDXL) while maintaining visual quality close to full-precision outputs.
-
Achieve this performance with minimal computational and memory overhead during inference, making these powerful generative capabilities viable for resource-constrained environments (like edge devices or real-time applications).
-
Provide a robust method for deploying quantized diffusion models in production pipelines where the primary challenge is mitigating the accumulation of quantization errors over time steps.
Sources
- PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
- DiTAS: Quantizing Diffusion Transformers via Enhanced Activation Smoothing
- IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models
- EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models
- BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
- HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
- EDA-DM: Enhanced Distribution Alignment for Post-Training Quantization of Diffusion Models
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- Efficient Diffusion Models: A Survey
- TMPQ-DM: Joint Timestep Reduction and Quantization Precision Selection for Efficient Diffusion Models
- QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning
- SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
- An Analysis on Quantizing Diffusion Transformers
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models