Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones
Ayoub Ghriss, Sourav Chakraborty
cs.LG, cs.AI, cs.IT, stat.ML
Submitted: 2026-07-19
Comments: 39 pages, 4 figures, 10 tables. Code: https://github.com/ayghri/kata
Code: https://github.com/ayghri/kata
License: http://creativecommons.org/licenses/by/4.0/
The gist: Linear attention promises constant-time recurrent inference but degrades sharply on associative recall.
Terminology
Abstract
Linear attention promises constant-time recurrent inference but degrades sharply on associative recall. We formulate attention recall as a spherical-packing problem and introduce Kernelized Linear Attention Activations (KATA), a framework whose feature maps are derived from first principles by certifying nonnegative attention weights through a self-dual homogeneous cone. Building on this observation, we show that rank-one positive semi-definite (PSD) features offer a favorable capacity--interference tradeoff. KATA recovers a parameter-free convex output gate and characterizes associative capacity through the Welch interference floor. For tolerances above this floor, KATA enlarges the state without adding parameters and admits spherical codes with exponentially many keys in the projection dimension. We implement KATA as fused Triton kernels at two operating points: a flash-attention-style forward up to about 1.6 times FlashAttention-2 throughput, and an exact O(T) chunked-state form that reaches about 11 times FlashAttention-2 forward throughput at 131 k tokens. An associative scan of the first-order feature lowers the inter-chunk recurrence depth to O((T/C)) for chunk size C and averages about 2.4 times the throughput of a matched sequential linear-attention baseline. On long-range MQAR and repeated-key overwrite, several KATA variants outperform Gated DeltaNet, with parameter counts and state sizes reported alongside accuracy. Induction preserves near-perfect recall, while kernel benchmarks show that the maps can be implemented efficiently. KATA retains 0.985 MQAR at a 16 times out-of-distribution length, approaching the softmax with roughly one quarter of the KV-cache entries. Experiments on 340M-parameter LLMs reveal a feature-dependent fluency trade-off and clarify how positional embeddings, delta rules, and decay gates interact with feature geometry.
Sources
- Zoology: Measuring and Improving Recall in Efficient Language Models
- Simple linear attention language models balance the recall-throughput tradeoff
- Titans: Learning to Memorize at Test Time
- Critical attention scaling in long-context transformers
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- MoM: Linear Sequence Modeling with Mixture-of-Memories
- Tables of the existence of equiangular tight frames
- Repeat After Me: Transformers are Better than State Space Models at Copying
- Modern Methods in Associative Memory
- LoLA: Low-Rank Linear Attention With Sparse Caching
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- Retentive Network: A Successor to Transformer for Large Language Models
- Gemma 3 Technical Report
- Qwen3 Technical Report
- Understanding Transformer from the Perspective of Associative Memory
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks