Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration
cs.LG, cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- Online Vector Quantized Attention
- Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
- Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
- RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
- Mistral 7B
- The Key to Going Linear: Analysis-Driven Transformer Linearization
- Liger: Linearizing Large Language Models to Gated Recurrent Structures
- Linearizing Large Language Models
- FP8 Formats for Deep Learning
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- Qwen2.5 Technical Report
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- AVQ-Attention: Adaptive Vector-Quantized Attention
- RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State
- The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
- LoLCATs: On Low-Rank Linearizing of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks