Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

arXiv:2608.13426 · cs.LG, cs.AI, cs.CL · Submitted 2026-08-13 · Read on arXiv

University of Chicago · Independent Researcher · Stony Brook University

cs.LG, cs.AI, cs.CL

Submitted: 2026-08-13

Updated: 2026-08-31

Comments: 24 pages

Code: https://github.com/Zesearch/rmm-llm

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

Terminology

Summary

Summary

This paper introduces Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights. Under a simple retention-ratio control, RMM provides a smooth and predictable accuracy–efficiency trade-off.

The core idea is to unify Transformer computations (attention and MLP layers) into the form Y = AB, where A is the activation matrix determined by the current input, B is a weight matrix or intermediate representation. RMM selects an index set I ⊆ [d] with I = ⌈ρd⌉ (where ρ is a user-controlled retention ratio) and computes RMMρ(A, B) ≜ A:,I BI,:. The selection is done by assigning each feature dimension j an importance score sj = ∥A:,j∥2 (the magnitude of feature j under the current input) and selecting I = TopK(sj dj=1, ⌈ρd⌉). This procedure is fully deterministic and adapts to each input, layer, attention head, and decoding step.

The paper provides theoretical justification: Theorem 1 proves that TopK selection by column norm is minimax optimal, minimizing the worst-case approximation error over all possible B at any given retention budget. Proposition 1 derives the approximation error bound: ∥AB − A:,IBI,:∥F ≤ Σj∈Ī ∥A:,j∥2∥Bj,:∥2, and Corollary 1 gives a factorized bound via Cauchy–Schwarz: ∥AB − A:,IBI,:∥F ≤ ∥A:,Ī∥F∥BĪ,:∥F.

RMM is applied to attention and MLP layers as follows:

  • Attention: For a single attention head with queries Q, keys K, and values V, RMM computes feature scores sj = ∥Q:,j∥2 and selects I = TopK(sj dj=1h, ⌈ρddh⌉), yielding reduced attention scores Se = (1/√dh)Q:,IK⊤:,I. Attention weights are then P = softmax(Se + M). Optionally, the attention–value multiplication P V is sparsified over the token dimension by computing token scores at = ∥P:,t∥2, selecting T = TopK(at Lkt=1, ⌈ρtLk⌉), and evaluating Oe = P:,TVT,:.

  • MLP and linear projections: Given activations X and weights W, RMM computes feature scores sj = ∥X:,j∥2, selects I = TopK(sj dj=1, ⌈ρdd⌉), and evaluates Ye = X:,IWI,:.

For complexity, a matrix multiplication A ∈ Rn×d and B ∈ Rd×m costs O(ndm) densely. With feature retention ratio ρd, RMM reduces the cost to O(nρddm). In attention, reducing QK⊤ over the head dimension lowers the cost from O(LqLkdh) to O(LqLkρddh). If token selection is applied to P V, the cost is reduced from O(LqLkdh) to O(LqρtLkdh). The overhead of computing feature scores (O(nd)) and top-k selection is small relative to the dense matrix multiplications.

Experimental setup: The paper evaluates RMM on a wide spectrum of pre-trained LLMs, including Llama 3.1 70B, Llama 3.1 8B, Llama 3.2 3B, Llama 3.2 1.5B, Qwen3 32B, Qwen 3.1 7B, and Qwen2.5-VL-7B-Instruct. Benchmarks span general QA and reasoning (COPA, PIQA, CommonsenseQA, ARC-Easy, ARC-Challenge, MMLU), language modeling (WikiText, BookCorpus), mathematics and coding (GSM8K, HumanEval), long-context reasoning (RULER-CWE, RULER-Hotpot), summarization (CNN/DailyMail), and vision–language tasks (POPE, Blink Art Style, Blink Forensic Detection, Blink Counting). All tasks are evaluated in the zero-shot setting. Baselines include static pruning methods (SparseGPT, Wanda, SliceGPT, magnitude pruning), dynamic baselines (H2O), a static variant of RMM, and random pruning.

Main results:

  1. Controlled comparison of pruning behavior (LLaMA 3.1 8B, RR = 0.5): RMM achieves the best average accuracy among all pruning methods on zero-shot QA benchmarks (59.8% average vs. 56.1% for SparseGPT, 52.7% for Wanda, 37.0% for SliceGPT, 39.3% for magnitude pruning). On abstractive summarization (CNN/DailyMail), RMM at RR = 0.8 remains close to the full model (ROUGE-1: 37.5 vs. 37.4 baseline), while at RR = 0.5 it still outperforms all baselines (ROUGE-1: 34.2 vs. 28.0 for static, 5.7 for random, 24.4 for H2O).

  2. Scaling across models and retention ratios: At matched retention levels, larger models generally retain stronger performance under moderate reduction. At RR = 0.8, LLaMA 3.1 70B remains close to the full model on most benchmarks, whereas smaller models show more pronounced degradation. As the retention ratio decreases, smaller models exhibit an earlier performance inflection (noticeable drops at RR = 0.7), while larger models degrade more gradually and maintain higher absolute accuracy even under aggressive pruning (RR ≤ 0.6).

  3. Stability under generation and long-context settings: RMM remains robust across generative and long-context settings. At RR = 0.8, outputs remain largely consistent with the full model; at RR = 0.6, generated text shows simplification and stylistic drift but remains coherent. On RULER long-context benchmarks, RMM achieves performance comparable to the full model across context lengths up to 30K tokens at both RR = 0.8 and RR = 0.5, with no systematic increase in degradation as context length grows.

  4. Vision-language generalization: On Qwen2.5-VL-7B, RMM at RR = 0.8 achieves performance nearly identical to the full model across all benchmarks (e.g., POPE: 82.0 vs. 83.7 baseline). At RR = 0.5, it still retains strong performance and substantially outperforms static and random pruning. Qualitative attention map visualizations show RMM preserves dense-like attention to relevant visual regions, while static and random pruning produce less aligned patterns.

  5. Ablation and mechanistic analysis: The paper validates two key design choices: dynamic selection (vs. static variant) and activation-aware scoring (vs. random pruning). Both are essential to RMM's effectiveness. Component-wise analysis reveals a clear structural asymmetry: attention-side computations are substantially more reducible than MLP components. Pruning attention-related operations degrades gradually, while pruning MLP components leads to much sharper performance drops. Within the MLP, the Up projection is most sensitive (16.32-point drop on ARC-Easy at RR = 0.7), the Gate projection intermediate (7.20-point drop), and the Down projection most robust (3.51-point drop). Pruning the entire MLP block causes severe performance collapse.

  6. Wall-clock latency: On NVIDIA A100 with LLaMA 3.1 8B (batch size 1, RR = 0.8), kernel-level benchmarks show speedups of 1.36×–1.67× for QK⊤ and 1.67×–1.89× for AV operations across sequence lengths 1024–4096. End-to-end latency speedups range from 1.05× at sequence length 1024 to 1.40× at 4096. For LLaMA 3.1 70B, speedups are 1.03× at 1024 and 1.41× at 2048, and RMM avoids out-of-memory at 4096 where the dense implementation fails.

Additional findings from appendices: RMM is competitive with or stronger than TEAL (an activation-sparsity method) in most comparisons, and extends reduction scope to internal attention products QK⊤ and PV. RMM transfers across multiple VLM architectures (LLaVA-1.5-7B, Gemma 3 12B, InternVL3-8B). RMM is compatible with INT8 weight quantization. Perplexity results show gradual degradation as RR decreases, with sharp degradation below 0.6, and larger models show greater robustness.

The paper concludes that redundancy in Transformer inference is not uniformly distributed across components, suggesting efficient inference methods should account for structural differences. RMM positions matrix-product-level adaptive reduction as a promising direction for efficient Transformer inference.

Improvements for AI systems

Improvements to AI Systems Based on RMM:

  1. Input-Adaptive Inference Engine: Implement a runtime layer that dynamically prunes matrix multiplications in attention and MLP blocks based on per-input activation magnitudes, reducing compute by 20–50% without retraining or weight modification. This enables faster inference on resource-constrained devices (e.g., mobile, edge) while maintaining near-full accuracy.

  2. Structured Redundancy-Aware Pruning Scheduler: Use RMM’s finding that attention components are more reducible than MLP components (and that Up/Gate/Down projections have differing sensitivity) to allocate retention ratios per layer and per submodule—e.g., keep MLP Up projections at ρ=0.9 while reducing attention QK T to ρ=0.5. This yields higher accuracy-per-FLOP than uniform pruning.

  3. Adaptive Long-Context Processing: For models handling sequences up to 30K+ tokens, apply RMM’s token-selection mechanism to attention-value multiplication, reducing memory and compute in KV-cache operations. This allows longer context windows on fixed hardware, improving performance on document summarization, code repositories, and multi-hop reasoning tasks without quality loss at ρ=0.8.

  4. Vision-Language Efficient Decoding: Integrate RMM into VLMs (e.g., Qwen2.5-VL, LLaVA) to prune cross-attention and MLP computations during image-grounded generation. This reduces latency for real-time applications like visual question answering, image captioning, and robotic perception, while preserving spatial attention to relevant visual regions.

  5. Quantization-Compatible Acceleration: Combine RMM with INT8 weight quantization to achieve compounded speedups (e.g., 1.5–2× from pruning + 2–4× from quantization) on GPUs and specialized hardware, enabling deployment of 70B-class models on single consumer GPUs for interactive use.

  6. Predictable Quality-Degradation Control: Expose a single retention-ratio hyperparameter (ρ) to users, allowing smooth, predictable trade-offs between speed and accuracy—e.g., ρ=0.8 for near-lossless performance, ρ=0.5 for 1.4× speedup with minor quality drop. This is ideal for adaptive systems that adjust compute based on battery level, network bandwidth, or real-time constraints.

  7. Training-Free Model Compression Pipeline: Provide a drop-in module that analyzes any pre-trained Transformer (LLM or VLM) and generates a pruned inference plan without fine-tuning. This reduces deployment time from days (for distillation or structured pruning) to minutes, and works across architectures (Llama, Qwen, Gemma, InternVL).

  8. Stable Generative Decoding: Use RMM’s demonstrated stability in long-form generation to reduce inference cost in chatbots, code completion, and creative writing tools, ensuring coherent outputs even at aggressive pruning (ρ=0.6), with only stylistic drift—enabling faster interactive responses.

  9. Hardware-Aware Kernel Optimization: Implement RMM’s top-k selection and sliced matrix multiplication as fused CUDA kernels (as benchmarked on A100), achieving 1.36–1.89× speedups on attention operations. This can be integrated into inference frameworks (e.g., vLLM, TensorRT-LLM) for immediate throughput gains in production serving.

  10. Redundancy Profiling for Model Design: Use RMM’s component-wise sensitivity analysis (attention vs. MLP, Up vs. Gate vs. Down) to inform future model architectures—e.g., designing models with inherently more prunable attention layers or allocating capacity to sensitive MLP projections, leading to more efficient base models.

Abstract

Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights. Under a simple retention-ratio control, RMM provides a smooth and predictable accuracy-efficiency trade-off. Across language models ranging from 1B to 70B parameters, we find that reduction tolerance depends on the model family, task, component, and retention ratio, although it often improves with model scale. Under moderate reduction, RMM remains robust across the evaluated discriminative, autoregressive generation, and long-context settings. We further show that the same principle extends to multimodal vision-language inference. Mechanistic ablations reveal a structural asymmetry within Transformers: attention-side computations are substantially more reducible than MLP components. Finally, wall-clock benchmarks with custom kernels on an NVIDIA A100 show that these computational savings can translate into practical runtime gains, especially at longer sequence lengths. Together, these results position RMM as a scalable direction for input-adaptive inference-time optimization.

Sources

Related papers