Attention Sinks and Outliers in Attention Residuals
cs.LG, cs.AI
Submitted: 2026-05-18
Updated: 2026-09-24
Code: https://github.com/kyegomez/attn_
Terminology
Sources
- Training Verifiers to Solve Math Word Problems
- GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
- The Llama 3 Herd of Models
- InnerQ: Hardware-Aware Tuning-Free Quantization of KV Cache for Large Language Models
- Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
- Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
- Attention Residuals
- Qwen3 Technical Report
- Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks