Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention
cs.LG, cs.AI
Submitted: 2026-08-28
Updated: 2026-09-22
Comments: 16 pages, 1 figure
License: http://creativecommons.org/licenses/by/4.0/
The gist: Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row- 1 approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued
Terminology
Abstract
Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row- 1 approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued output. Two sharp worst-case laws isolate support geometry: for fixed d and error epsilon, spherical self-attention has rank Θ d, epsilon(n,(1+β)(d-1)/2), while full-ball geometry adds one radial degree and, for β β 0(d, epsilon) and n C d e β/8, gives Θ d, epsilon(β d/2). For a fixed head, row-softmax quotients out row-scalar logit directions: the remaining visible query--key interaction dimension r yields an r/2 per-instance upper law, and bounded constructions show this exponent is minimax sharp. Approximate interaction subspaces incur an explicit residual output error and yield a tolerance-indexed SVD dimension. On an 84-head BERT-base calibration set, we observe modest effective-dimension reductions across many head--temperature settings, together with positive associations with finite constructive rank upper certificates. Together, these results separate support geometry, which sets worst-case temperature scaling, from softmax-visible interaction geometry, which controls per-head approximation complexity.
Sources
- Optimal Volume-Sensitive Bounds for Polytope Approximation
- Economical Convex Coverings and Applications
- Hyperbolic contractivity and the Hilbert metric on probability measures
- Compressible Softmax-Attended Language under Incompressible Attention
- On the Invariants of Softmax Attention
- Nearly Optimal Attention Coresets
- Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
- The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks