Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling
summary
The gist
Quantization noise can accumulate over diffusion model denoising trajectories, degrading generation quality, and this paper introduces Q-Drift, a sampler-side correction that treats quantization
In short
Q-Drift corrects generation quality lost due to quantization noise accumulating during diffusion model sampling. It treats quantization error as a stochastic perturbation, deriving a drift adjustment that preserves the marginal distribution of the generated output. This sampler-side correction improves metrics like FID without significant inference overhead.
Key concepts
- Quantization Noise Accumulation
- When large diffusion models are deployed using post-training quantization, small errors occur at each step. These errors do not cancel out; instead, they accumulate along the entire denoising trajectory. This accumulation degrades the final image quality because the model deviates from the true data distribution.
- Implicit Stochastic Perturbation
- The paper models quantization error as an invisible noise injected into every denoising step, similar to a stochastic process. Instead of treating it as a fixed error, Q-Drift treats this noise implicitly during sampling to derive a correction that accounts for its cumulative effect on the distribution.
- Marginal-Distribution-Preserving Drift Adjustment
- The goal is not just to fix errors at one step but to ensure the overall distribution of generated samples remains accurate. Q-Drift calculates a specific adjustment factor based on this noise model, which scales the sampler's movement to counteract the drift caused by quantization error.
- Q-Drift Update Formula
- The final correction involves scaling the deterministic update step: x(n+1) = x(n) + Δσ_i * (1 + c_i)⊙ ϵˆ. This formula incorporates a correction factor (c_i) derived from the noise variance, effectively adding a controlled stochastic component to the standard sampling process.
Terminology used across episodes
This episode discusses
- Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling · Paper Radio
- PixArt-: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
- DiTAS: Quantizing Diffusion Transformers via Enhanced Activation Smoothing
- IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models
- EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models
- BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
- HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
- EDA-DM: Enhanced Distribution Alignment for Post-Training Quantization of Diffusion Models
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- Efficient Diffusion Models: A Survey
- TMPQ-DM: Joint Timestep Reduction and Quantization Precision Selection for Efficient Diffusion Models
- QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning
- SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
- An Analysis on Quantizing Diffusion Transformers
The paper
Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling · Read on arXiv
Department of Computer Science and Engineering, Seoul National University · School of Computer and Communication Sciences, EPFL · Meta Reality Labs
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling".
Jane: Quantization noise can accumulate over diffusion model denoising trajectories, degrading generation quality, and this paper introduces Q-Drift,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's start by looking at the title and who came up with this work. It’s "Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling," and it was written by Sooyoung Ryu, Mathieu Salzmann, and Saqib Javed. The name itself suggests they are focusing on correcting a drift issue specifically related to quantization in diffusion sampling methods.
Jane: That title tells us exactly what the core problem is: we're dealing with a specific type of drift that happens when you quantize models used in diffusion sampling, so it’s very targeted research. It’s about keeping the sampling process aligned with what a full-precision model would do, even under compression.
Lu: The authors are coming from institutions like Seoul National University and EPFL, which suggests they bring a strong theoretical background to this problem of modeling stochastic perturbations in continuous time dynamics, which is where diffusion sampling lives.
Meng: I've seen work on quantization stabilization before, but what makes this paper different is its focus on the trajectory drift rather than just stabilizing the output at a single step; that's a crucial distinction for deployment.
Lalam: I think it’s important to hear about their approach because if they can provide a principled way to adjust the drift, it gives us more confidence in pushing these models into real-world applications where storage and inference speed are critical constraints.
The paper's summary: Tom: So, what is the main gist of what "Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling" actually proposes? Essentially, they introduce Q-Drift as a sampler-side correction that treats the quantization error as an implicit stochastic perturbation at every denoising step to derive an adjustment that keeps the marginal distribution preserved.
Jane: In simpler terms, it means instead of just fixing the image quality at one point in time, they are modeling how that noise evolves over the entire sequence of steps, which is what causes the degradation we see in practice. They derive a specific formula for this drift adjustment based on empirical observations from models like D2-DPM.
Lu: The methodology involves modeling the quantization error as an implicit Gaussian perturbation injected into the denoiser output at each step, then using a generalized SDE framework to reparameterize the marginal-preserving family in terms of a noise level sigma(t).
Meng: That sounds mathematically intensive; I need to know if this is something that can be implemented without requiring an overhaul of the core sampling loop or adding significant computational overhead during inference.
Lalam: The paper summarizes that they estimate a timestep-wise variance statistic from calibration, and they show that this correction factor is plug-and-play with common samplers and PTQ methods, which makes it very appealing for immediate practical use.
The paper's improvements: Tom: Moving on to the specific improvements the authors detail, they focus on deriving a correction factor, c i, based on matching an implicit stochastic increment variance to the diffusion term in the Euler–Maruyama discretization for a first-order sampler.
Jane: They show that this leads to a final Q-Drift update formula where the next state is calculated by scaling the deterministic Euler update component with (one + c i), which explicitly incorporates that derived correction factor into the sampling process.
Lu: The derivation involves matching this implicit variance to the diffusion term, resulting in c i = sigma i / (two sigma i) V sigma i, where V sigma i is estimated during an offline calibration phase. This shows they are deriving a statistically grounded correction rather than just guessing what might work.
Meng: The calibration requirement is interesting; they state that this variance statistic can be estimated with as few as five paired full-precision and quantized calibration runs, which makes the setup incredibly practical for deployment environments where extensive data collection is not feasible.
Lalam: And they verified that these estimated correction factors stay stable even when you use nested subsamples of a standard 5K-prompt calibration run, which speaks to the robustness of their statistical estimation process.
Conclusion: Tom: So, to wrap up what we've discussed about "Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling," the main point is that this method provides a principled way to correct for the drift caused by accumulating quantization noise during diffusion sampling trajectories.
Jane: They show that by treating the quantization error as an implicit perturbation, we can derive a marginal-distribution-preserving adjustment that works across various samplers and architectures without needing to change the underlying model or quantization scheme.
Lu: The paper’s extension of this drift-rescaling principle to other samplers, like Flow-matching and DPM-Solver++, is quite telling; it suggests a general mathematical principle governing how we can handle this specific type of error correction in diffusion dynamics.
Meng: From an engineering perspective, the negligible overhead at inference mentioned by the authors is what makes me most optimistic; if we can get this kind of fidelity improvement without adding significant latency, it becomes much more viable for real-time applications.
Lalam: I think the ultimate implication here is that we can deploy these large diffusion models in production pipelines where the primary challenge was previously treating quantization error as purely local, leading to a consistent visual quality improvement across diverse models.
More episodes
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck
- 2407.14562-Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Bootstrapped Prolog-based Chain-of-Thought