Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
cs.LG, cs.CL
Submitted: 2026-09-18
Updated: 2026-09-29
Comments: 21 pages (8 pages main text plus appendices), 5 figures, 12 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Capacity-Based Rationale for Multi-Head Attention
- On the Complexity of Neural Computation in Superposition
- Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
- Why do LLMs attend to the first token?
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
- Near-Floor Geometry Is Generic: Leverage Dispersion in Trained Overcomplete Codes
- Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
- Vision Transformers Need Registers
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- A Unifying View of Attention Sinks: Two Algorithms, Two Solutions
- Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
- IMU-1: Sample-Efficient Pre-training of Small Language Models
- When Attention Sink Emerges in Language Models: An Empirical View
- Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
- Mathematical Models of Computation in Superposition
- Do Value Vectors in Deep Layers Need Context from the Residual Stream?
- Transformer Quality in Linear Time
- The Routing and Filtering Structure of Attention
- Vision Transformers Don't Need Trained Registers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks