When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation
Shuhan Wang, Yilin Luo, Nan Xu, Chi Wang Cheung
cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/zhangir-azerbayev/proof-pile
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 81/100
The gist: Based on the paper, here is a detailed summary: The paper investigates whether rotation-based post-training quantisation (PTQ) can be improved by respecting the structural decomposition imposed by
Terminology
Summary
Based on the paper, here is a detailed summary:
The paper investigates whether rotation-based post-training quantisation (PTQ) can be improved by respecting the structural decomposition imposed by Rotary Position Embedding (RoPE). Specifically, it asks whether a rotation that operates within each two-dimensional RoPE frequency pair (pairwise mixing) can outperform a full-head Hadamard transform that mixes all channels in an attention head, under dynamic W4A4KV4 quantisation.
Theoretical Contributions:
The paper establishes a converse characterisation of the RoPE centraliser. For a single attention head with distinct RoPE frequencies, the only orthogonal maps that commute with RoPE are independent planar rotations within each frequency pair (Lemma 2). This family was previously used by FPTQuant (van Breugel et al., 2025), but the paper proves that no other single-head orthogonal map commutes with RoPE. For the head-shared parameterisation (where one angle is shared across all heads for each layer and frequency pair), the paper derives a closed-form angle (Theorem 3) that minimises the larger channel variance under a pooled-covariance, position-averaged surrogate. The implementation is verified to attain this analytic minimum with a maximum excess of 6.90 × 10−7.
Empirical Findings (Negative Result):
The core empirical finding is a scoped negative result. Across four checkpoints (Llama-3.2-1B, Llama-3.2-3B, Llama-3.1-8B, Mistral-7B-v0.3), replacing the full-head Hadamard with the head-shared pairwise configuration increases perplexity at both short and long context lengths. For example, on Llama-3.2-1B the difference is +1.33 PPL, and on Mistral-7B-v0.3 it is +0.05 PPL. The paper states: Across the evaluated Llama and Mistral checkpoints, replacing the full-head Hadamard with the head-shared pairwise configuration increases perplexity in every short-context comparison.
This ordering holds at every evaluated long-context point, including 128K where available. Composing the pairwise rotation with the Hadamard does satisfy the selected ±0.05-PPL interval criterion under the default estimator.
Analysis of the Discrepancy:
The paper examines why the optimised local surrogate does not translate to better quantisation performance. Two factors are identified:
-
Estimator mismatch: The default angle estimator pools Q and K calibration statistics, even though only K is quantised after the R3 transform. Estimating the shared angle from K alone improves the pairwise-only configuration on all four checkpoints, but does not eliminate the gap relative to full-head mixing.
-
Objective and support mismatch: The analytic objective controls a position-averaged second moment within each pair, whereas the dynamic INT4 quantiser determines its step from a tokenwise group range. The pairwise transform also provides only two-channel support for redistributing a peak, compared to the full-head Hadamard's dh-channel support. A controlled interpolation from two-channel to full-head mixing shows that K range, relative quantisation error, and perplexity degradation all decrease as support increases.
Conclusion:
The paper concludes that optimality for a structured surrogate need not reduce quantisation error when the surrogate and mixing support are misaligned with the quantiser’s scale-setting statistic.
The broader implication is that structural alignment alone does not determine whether a rotation reduces quantisation error; the transform must be judged by whether its optimisation objective and mixing support match the scale-setting rule of the quantiser in which it is deployed.
Improvements for AI systems
Improvements to AI Systems:
-
Quantisation-Aware Rotation Selection: Instead of using a single fixed rotation (e.g., Hadamard) for post-training quantisation, the AI system can dynamically select or compose rotations based on the quantiser’s scale-setting statistic (e.g., tokenwise group range). The system will evaluate candidate rotations (pairwise, full-head, or hybrid) against the actual quantiser’s error metric, not just a variance surrogate, and choose the one that minimises perplexity or downstream task loss.
-
Support-Matched Mixing for Outlier Redistribution: The system can adapt the mixing support (number of channels involved in a rotation) to match the quantiser’s group size. For example, if the quantiser uses a group size of 128 channels, the rotation will be designed to spread outliers across exactly that many channels, rather than using a fixed 2-channel (pairwise) or full-head support. This reduces relative quantisation error and perplexity degradation by aligning the rotation’s redistribution capability with the quantiser’s step-size computation.
-
Separate Q/K Rotation Optimisation: The system will estimate and apply distinct rotations for query (Q) and key (K) projections, since only K is quantised in many R3-based pipelines. By optimising the rotation angle from K’s calibration statistics alone (not pooled Q/K), the system reduces quantisation error and improves perplexity on pairwise-only configurations, as demonstrated in the paper’s analysis.
-
Surrogate-Objective Alignment Check: Before deploying a rotation, the AI system will run a quick diagnostic to verify that the rotation’s optimisation objective (e.g., minimising channel variance) actually correlates with the quantiser’s error. If the correlation is weak (as found with the position-averaged surrogate), the system will switch to a different rotation or quantisation scheme, avoiding the false confidence from theoretically optimal but practically ineffective transforms.
-
Adaptive Rotation Composition for Long-Context Robustness: The system will monitor perplexity at long context lengths (e.g., 128K) and automatically compose a pairwise rotation with a full-head Hadamard if the pairwise-only configuration degrades performance. This ensures that the chosen rotation remains robust across varying sequence lengths, as the paper shows the pairwise configuration consistently underperforms at all long-context points.
What the Improved AI System Can Do:
-
Achieve lower perplexity and higher task accuracy under aggressive W4A4KV4 quantisation by selecting rotations that match the quantiser’s scale-setting rule.
-
Automatically adapt its rotation strategy to different model architectures (e.g., Llama vs. Mistral) and quantisation hardware constraints (e.g., group size, bit-width) without manual tuning.
-
Provide a principled, quantiser-aware post-training quantisation pipeline that avoids the pitfall of optimising for a surrogate that does not translate to real quantisation error reduction.
-
Maintain performance at both short and long context lengths, ensuring reliability for applications like retrieval-augmented generation or long-document processing.
Abstract
Rotation-based post-training quantisation commonly applies an orthogonal transform across an entire attention head to reduce outlier-induced error. RoPE instead partitions each head into two-dimensional frequency pairs, raising the question of whether a transform respecting this decomposition can improve on full-head mixing. Prior work has established the per-pair rotations that commute with RoPE. We state the converse result that, for distinct frequencies, no other single-head orthogonal map commutes with RoPE. For the head-shared parameterisation used in our experiments, we then derive the rotation angle that minimises the larger channel variance under a pooled-covariance, position-averaged surrogate and verify that the implementation attains its analytic minimum. The evaluated head-shared pairwise configuration does not improve accuracy in the tested dynamic W4A4KV4 setting. Across four checkpoints, replacing the full-head Hadamard with this configuration increases perplexity at both short and long context lengths. Composing the pairwise rotation with the Hadamard satisfies the selected plus or minus0.05-PPL interval criterion under the default estimator. Estimating the shared angle from K alone improves pairwise-only on every checkpoint but does not close its gap to full-head mixing. The analytic objective controls a position-averaged second moment of a pooled calibration covariance, whereas the dynamic quantiser sets its step from a tokenwise group range. The pairwise transform also has only two-channel mixing support. Along a controlled interpolation from two-channel to full-head mixing, K range, relative quantisation error, and perplexity degradation decrease as support increases. These results show that optimality for a structured surrogate need not reduce quantisation error when the surrogate and mixing support are misaligned with the quantiser's scale-setting statistic.
Sources
- Intriguing Properties of Quantization at Scale
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- KurTail : Kurtosis-based LLM Quantization
- SliceGPT: Compress Large Language Models by Deleting Rows and Columns
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- QuIP: 2-Bit Quantization of Large Language Models With Guarantees
- WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
- PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
- Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- The Llama 3 Herd of Models
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
- Efficient Attentions for Long Document Summarization
- SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
- Mistral 7B
- CommVQ: Commutative Vector Quantization for KV Cache Compression
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks