Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages

arXiv:2608.11786 · cs.CL · Submitted 2026-08-12 · Read on arXiv

Nirmal Thomas

Prathama International

cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 9 pages, 1 figure, 6 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Language-Conditional Dequantization (LCD) is a post-hoc method that attaches per-language rank-2 LoRA corrections to the linear layers of an already-quantized model, adding 0.12% parameters per

Terminology

Summary

Language-Conditional Dequantization (LCD) is a post-hoc method that attaches per-language rank-2 LoRA corrections to the linear layers of an already-quantized model, adding 0.12% parameters per language and training in under 20 minutes on a single GPU. Across Qwen2.5-3B and Llama-3.2-3B, LCD recovers 70–83% of the perplexity gap for non-Latin-script languages and 17–28% of the GlobalMMLU accuracy gap, outperforming a language-agnostic correction of equal capacity by 3–9 points on typologically distant languages and a data-free low-rank baseline (LQER) by an order of magnitude. We further identify a perplexity–accuracy disconnect and trace it to where quantization concentrates damage: early-depth errors (Llama) propagate downstream and resist local correction, while late-depth errors (Qwen) do not. A layer-restricted variant of LCD validates this mechanism directly.

Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2–4× larger perplexity degradation on non-English languages than on English. On Qwen2.5-3B, INT3 GPTQ degrades Arabic perplexity by 4.37× and Japanese by 3.04×, compared to just 1.35× for English; on Llama-3.2-3B the pattern repeats (Arabic 3.65×, English 1.39×). In the sub-4B INT3 regime, this is a practically consequential, language-dependent quality gap.

Multilingual calibration addresses the root cause but requires re-quantization, which is impractical for already-deployed models. Static error correction methods (LQER, RILQ, ResQ) add low-rank residuals post-hoc but apply the same correction regardless of input language. Input-conditional methods like BinaryMoS adapt weights to the input but target binarization and ignore language identity. LCD is built on a simple insight: if quantization error is language-dependent, the correction should be too. For each linear layer in a quantized model, LCD attaches a per-language rank-r additive correction via forward hooks: ŷ = Wq x + (Al · Bl) · x, where Al ∈ Rdout×r and Bl ∈ Rr×din are language-specific correction matrices. With rank r = 2, this adds only 0.12% parameters per language. Corrections are trained on 256 samples of language-specific text in under 20 minutes per language on a single GPU. At inference time, the appropriate correction is selected based on the input language, requiring no modification to the base model architecture.

The contributions are: (1) characterizing the multilingual degradation caused by English-calibrated INT3 GPTQ on two sub-4B models, quantifying per-language perplexity ratios from 1.35× (English) to 4.37× (Arabic); (2) proposing LCD, a post-hoc method that recovers 70–83% of the perplexity gap for non-Latin-script languages at 0.12% parameter cost per language, without re-quantization; (3) showing that per-language conditioning captures real language-specific signal, as a language-agnostic rank-2 baseline matches per-LCD on average but trails by 3–9 points on typologically distant languages; (4) revealing a perplexity–accuracy disconnect traced to a structural cause: quantization concentrates damage in different network depths across models (Qwen: late; Llama: early-middle), validated by a layer-restricted variant of LCD.

For each linear layer and each language, the correction is defined as ycorrected = Wq x + Al Bl x, where Al and Bl are learnable parameters and 1/r is a fixed scaling constant. Al is initialized to zero and Bl ∼ N(0, 0.01), ensuring the correction is exactly zero at initialization. This follows the LoRA parameterization but serves a fundamentally different purpose: correcting systematic quantization error for a specific language distribution rather than adapting the model to a new task. Corrections are implemented as forward hooks on each linear layer (excluding lm head and embed tokens), requiring no modification to the model architecture or forward pass.

Each language is trained independently. Given a set of N = 256 text samples from the target language (sourced from mC4/C4), the standard language modeling loss is minimized with only the correction parameters trainable. AdamW is used with learning rate 5×10−4, cosine decay with 10% warmup, weight decay 0.01, gradient clipping at 1.0, and 500 steps per language. All base model parameters remain frozen. Training takes approximately 15–20 minutes per language on a single L4 GPU, so covering all eight non-English languages costs under 3 GPU-hours in total. Each trained correction ships as a ∼7 MB delta on top of the unchanged deployed checkpoint.

At inference time, the input language is identified (via locale, automatic detection, or explicit specification) and the corresponding correction slot is activated via a single index. When language is not known, off-the-shelf language identification (e.g., fastText LID) classifies a prompt in well under a millisecond. The base quantized model resides in memory once; language switching requires only updating the active adapter index, with no model reloading. For K languages with rank r = 2, total overhead is K × 0.12% additional parameters.

The experimental setup uses Qwen2.5-3B and Llama-3.2-3B, two sub-4B families with distinct tokenizers and training distributions. Quantization is GPTQ W3A16, group size 128, calibrated on 128 English C4 samples, replicating standard English-only deployment practice. Nine languages are evaluated: English (baseline), Arabic, Japanese, Chinese, Hindi, Russian, French, Spanish, Korean. Metrics are perplexity on held-out mC4/C4 (32 samples/language) and GlobalMMLU accuracy via log-likelihood over ≈14,040 items per language (57 subjects). Baselines are FP16 (unquantized), INT3 (uncorrected), and INT3+LCD (corrected). English is the calibration language and excluded from correction training.

On Qwen2.5-3B, English degrades by 1.35×, while Arabic degrades by 4.37×, Korean by 3.49×, and Japanese by 3.04× (a relative disparity of 2.3–3.2×). The pattern is consistent on Llama-3.2-3B, confirming the effect is not model-specific. Languages with non-Latin scripts and greater typological distance from English exhibit the largest degradation; Romance languages (French, Spanish) suffer only modest additional harm beyond English.

LCD recovers 70–83% of the quantization gap for non-Latin-script languages (Arabic, Japanese, Chinese, Korean) on both models. Recovery is lower for Latin-script languages close to English (French 38%/54%, Spanish 36%/52%), consistent with these languages suffering less language-specific quantization error. These figures are robust to random seed: retraining all corrections under three seeds changes per-language recovery by at most ±1.3 percentage points.

Comparing per-language vs. language-agnostic correction, a single shared rank-2 LoRA trained on mixed multilingual data matches per-LCD on average (65.4% vs. 64.6%), but the per-language breakdown reveals that per-language correction wins by 3–9 points on Arabic, Japanese, and Korean (precisely the languages whose distribution is most distant from the English calibration set) and loses 16 points on French, where no distinct per-language signal exists. This confirms that per-language conditioning captures genuine language-specific signal where such signal exists.

Against LQER, a data-free, untrained, language-agnostic baseline that reconstructs quantization error with rank-r truncated SVD of Wfp16 − Wq at each layer, at the identical rank-2 budget on Qwen2.5-3B, LQER recovers only 5.5% of the perplexity gap and 13.0% of the GlobalMMLU gap (non-English averages), versus 63.5% and 27.8% for LCD; LQER is even negative on several languages (e.g., Russian, −40% perplexity). A static SVD of the weight error cannot capture the input-dependent, language-specific component that LCD learns from data.

Recovery correlates with typological distance from English: non-Latin-script languages (Arabic, Japanese, Chinese, Korean) cluster at the top (70–83%), while Latin-script languages close to English (French, Spanish) cluster at the bottom (36–54%), with Hindi and Russian in between. This ordering is consistent across both models and confirms that LCD corrects calibration–distribution mismatch.

A rank ablation on Qwen2.5-3B at ranks 1, 2, and 4 (Arabic and Japanese) shows average recovery of 76.92% (r=1), 76.98% (r=2), and 76.84% (r=4). Rank is not the bottleneck: the gap from rank 1 to rank 4 is within 0.1 points, indicating the correction subspace is effectively one-dimensional on these languages. Rank r = 2 is adopted as a small safety margin; a rank-1 deployment would halve parameter overhead to 0.06% per language without measurable degradation.

On GlobalMMLU, INT3 quantization severely degrades accuracy on both models (Qwen: 64.7→41.8% English; Llama: 55.5→39.9%). On Qwen2.5-3B, LCD delivers consistent accuracy gains across every non-English language, with the largest improvements being Chinese (+9.1 points), Spanish (+6.7), Russian (+6.2), and Japanese (+6.0); Arabic, French, and Korean also gain between 3.5 and 7.8 points. Across the eight non-English languages, LCD closes an average of 28% of the FP16→INT3 accuracy gap. On Llama-3.2-3B, LCD also improves every non-English language, but gains are smaller, between 1.1 points (Chinese, Korean) and 3.5 points (Russian), closing an average of only 17% of the gap. This perplexity–accuracy disconnect is the paper’s most scientifically informative finding: Llama’s perplexity recovery (avg 69%) exceeds Qwen’s (avg 63%), yet its MMLU recovery lags substantially.

To diagnose the disconnect, per-layer relative Frobenius error ∥yfp16 − yint3∥F /∥yfp16∥F is measured for every linear layer’s output activations, separately for each language. Two structural findings explain the disconnect. First, the highest-error layers in both models are o proj and down proj, the compression bottlenecks that project from a larger intermediate space back to the hidden dimension, where GPTQ’s per-channel grid is least faithful. Second, the depth of these worst-error layers differs sharply: Qwen2.5-3B’s top-10 lie in indices 18–30 of 36 (late); Llama-3.2-3B’s lie in indices 7–14 of 28 (early-middle), with 14+ downstream layers consuming their corrupted output. Late-layer errors (Qwen) primarily distort final next-token logits, which perplexity measures and LCD fixes locally. Early-middle errors (Llama) distort intermediate representations that feed into downstream attention and MLP operations, exactly the multi-step computation MMLU requires. A rank-2 hook corrects a layer’s immediate output but cannot undo propagated upstream corruption.

Direct validation via layer-restricted LCD: on Llama, BOTTOM-HALF (covering the worst-error band at indices 7–14) exceeds uniform ALL by 10 pp and beats TOP-HALF by 43 pp on average, obtaining better recovery at half the parameters. On Qwen, the two halves are indistinguishable and both trail ALL, consistent with its worst-error layers spanning a broad 40% band. Critically, this perplexity advantage does not reach downstream accuracy: evaluated on GlobalMMLU, Llama’s BOTTOM-HALF LCD recovers only 9.5% of the accuracy gap, versus 17% for uniform ALL (the reverse of the perplexity ordering). Concentrating capacity where the distributional error is largest (the early compression band) actively trades away representational recovery. This double dissociation (bottom-half wins on perplexity but loses on MMLU) is the strongest evidence that the two levels of quantization damage are governed by distinct mechanisms, and that perplexity recovery is an unreliable proxy for downstream gains.

A negative result on narrow layer targeting: three configurations are evaluated — UNIFORM-R2 (rank-2, all layers), TARGETED-R4 (rank-4, layers 7–14 only), and TARGETED-R2 (rank-2, layers 7–14 only). UNIFORM-R2 achieves 45.0% average recovery; both targeted variants reach only 37%, and TARGETED-R4≈TARGETED-R2 (37.0% vs. 37.1%) rules out rank as an explanatory factor. The bottom-half advantage thus comes from broad coverage of the early depth band (layers 0–13 as a whole) rather than precise targeting of the identified hotspot. When correction capacity is limited, broad early-depth coverage dominates narrow layer allocation.

Limitations include: language identification at inference (code-switched inputs, low-quality language ID, or mixed-language prompts are not directly handled); residual MMLU gap on early-error architectures (allocating higher rank specifically to compression projections remains open); narrow quantization regime (only INT3 GPTQ W3A16, group size 128, English C4 calibration is studied); model scale and family coverage (at 7B, rank-2 LCD recovers only 8.5% of the non-English accuracy gap vs. 28% at 3B, and raising capacity to rank-4 with 1000 steps makes every language worse, as corrections overfit the 256-sample training slice); language coverage and data (only eight relatively high-resource non-English languages are evaluated); adapter training data (256 samples per language is deliberately small); and ethical considerations (adapters are trained on web-scraped text and may inherit biases; no audit of toxicity, factuality, or stereotype changes is performed).

The conclusion states that English-calibrated INT3 quantization creates a systematic, language-dependent quality gap that disproportionately harms non-English users. LCD demonstrates that this gap is largely correctable: rank-2 corrections at 0.12% parameter cost per language recover 70–83% of the perplexity degradation for non-Latin-script languages across two model families, and 17–28% of the GlobalMMLU accuracy gap. A language-agnostic baseline with the same capacity matches per-LCD on average but trails by 3–9 points on the typologically distant languages where language-specific signal is concentrated. The analysis further reveals that perplexity recovery does not reliably translate to downstream accuracy: Llama’s higher perplexity recovery yields smaller MMLU gains than Qwen’s, traced to depth asymmetry in where quantization concentrates damage. Restricting LCD to Llama’s bottom half confirms this (+10 pp over uniform), yet precision-targeting of layers 7–14 specifically does not help further (37% vs. 45%), establishing that broad early-depth coverage, not narrow layer allocation, is the operative mechanism. The broader implication is practical: quantization need not discriminate, and with trivial overhead, deployed quantized models can serve non-English users more equitably.

Improvements for AI systems

Improvements to AI Systems:

  1. Language-Aware Post-Hoc Quantization Correction: Integrate LCD as a plug-in module into any deployed quantized LLM (e.g., INT3/INT4 GPTQ). The system automatically detects input language (via fastText or locale) and applies per-language rank-2 LoRA corrections to linear layers, recovering 70–83% of perplexity degradation for non-Latin-script languages (Arabic, Japanese, Chinese, Korean) and 17–28% of GlobalMMLU accuracy gap, with only 0.12% parameter overhead per language and no re-quantization or architecture changes.

  2. Depth-Aware Correction Allocation: For models with early-layer quantization damage (e.g., Llama-3.2-3B), automatically allocate correction capacity to the bottom half of layers (indices 0–13) rather than uniformly. This yields +10 percentage points (pp) perplexity recovery over uniform allocation at half the parameters. For models with late-layer damage (e.g., Qwen2.5-3B), use uniform allocation across all layers. The system can detect the damage profile by measuring per-layer Frobenius error on a small calibration set and then choose the optimal layer-restricted strategy.

  3. Perplexity–Accuracy Disconnect-Aware Evaluation and Tuning: When optimizing for downstream tasks (e.g., MMLU), do not rely solely on perplexity recovery. The improved system uses a dual-metric validation: it tracks both perplexity and task accuracy during correction training. If perplexity improves but accuracy does not (as seen with Llama), the system automatically shifts correction capacity from early compression layers (o proj, down proj) to later layers or increases rank on those specific projections, preventing the trade-away of representational recovery.

  4. Typological-Distance-Adaptive Correction Strength: Automatically scale correction rank or training steps based on the typological distance of the target language from the calibration language (English). For non-Latin-script, distant languages (Arabic, Japanese, Korean), use rank-2 with full 500 steps; for close Latin-script languages (French, Spanish), use rank-1 with fewer steps (e.g., 200), halving overhead without measurable loss. This is justified by the finding that rank-1 achieves 76.92% recovery vs. 76.98% for rank-2 on distant languages.

  5. Zero-Shot Language Correction for Unseen Languages: Train a meta-correction model that predicts per-language LoRA corrections from language embeddings (e.g., from a multilingual sentence encoder) without any fine-tuning. This extends LCD to low-resource or unseen languages by interpolating corrections from typologically similar trained languages (e.g., Arabic correction for Urdu, Japanese for Korean), leveraging the observed correlation between typological distance and correction efficacy.

  6. Dynamic Language Switching with Shared Memory: Deploy a system where the base quantized model resides in GPU memory once, and language-specific corrections are stored as 7 MB deltas. The inference engine switches active adapters via a single index (sub-millisecond) without model reloading, enabling real-time multilingual serving for code-switched or mixed-language prompts by applying a weighted blend of corrections based on language identification confidence scores.

  7. Early-Layer Error Propagation Mitigation: For models with early-depth quantization damage, augment LCD with a lightweight representation repair pass: after applying layer-specific corrections, run a single forward pass with a learned residual connection that re-projects corrupted intermediate activations back toward the FP16 distribution. This addresses the finding that early errors propagate downstream and resist local correction, potentially recovering more than the current 17% MMLU gap for Llama.

  8. Adaptive Rank Selection via Validation Loss: During correction training, monitor validation perplexity on held-out language data. If rank-1 achieves within 0.1 pp of rank-2 (as observed), automatically reduce to rank-1, halving parameter overhead. If validation loss plateaus early (e.g., before 500 steps), early-stop to save compute. This makes the system cost-adaptive per language.

  9. Quantization-Regime-Aware Correction: Extend LCD to other quantization schemes (e.g., INT4, AWQ, or group sizes other than 128) by re-training corrections on the specific quantized model. The system includes a calibration step that measures per-language degradation ratios (e.g., 1.35× for English vs. 4.37× for Arabic) and automatically triggers correction only for languages exceeding a user-defined threshold (e.g., >2× degradation), avoiding unnecessary overhead for minimally affected languages.

  10. Bias and Safety Monitoring for Corrected Outputs: Since corrections are trained on web-scraped text, integrate an automated audit that compares corrected vs. uncorrected outputs for toxicity, factuality, and stereotype changes (using existing classifiers). If a language-specific correction introduces harmful shifts, the system falls back to the uncorrected quantized model for that language or applies a safety-aligned correction trained on filtered data.

What the Improved AI System Can Do:

  • Serve non-English users of quantized sub-4B models with near-FP16 quality for perplexity (70–83% gap recovery) and significantly better downstream accuracy (17–28% MMLU gap closure), at negligible cost (0.12% params, <20 min training per language).

  • Automatically adapt correction strategy to the model's damage profile (early vs. late layers), maximizing recovery without manual tuning.

  • Handle code-switched or mixed-language prompts by blending corrections in real-time.

  • Extend to unseen languages via typological interpolation, enabling equitable deployment for low-resource languages without additional training.

  • Avoid the perplexity–accuracy trap by optimizing for task-specific metrics, ensuring that corrections improve real-world performance, not just language modeling loss.

  • Maintain safety and fairness by auditing and filtering corrections, preventing amplification of web-scraped biases.

Abstract

Aggressive quantization disproportionately harms multilingual capability: in the sub-4B INT3 GPTQ regime, we measure 2-4x larger perplexity degradation on non-English languages than on English. We propose Language-Conditional Dequantization (LCD), a post-hoc method that attaches per-language rank-2 LoRA corrections to the linear layers of an already-quantized model, adding 0.12% parameters per language and training in under 20 minutes on a single GPU. Across Qwen2.5-3B and Llama-3.2-3B, LCD recovers 70-83% of the perplexity gap for non-Latin script languages and 17-28% of the GlobalMMLU accuracy gap, outperforming a language-agnostic correction of equal capacity by 3-9 points on typologically distant languages and a data-free low-rank baseline (LQER) by an order of magnitude. We further identify a perplexity-accuracy disconnect and trace it to where quantization concentrates damage: early-depth errors (Llama) propagate downstream and resist local correction, while late-depth errors (Qwen) do not. A layer-restricted variant of LCD validates this mechanism directly.

Sources

Related papers