Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention

arXiv:2608.28150 · cs.LG, cs.AI · Submitted 2026-08-28 · Read on arXiv

cs.LG, cs.AI

Submitted: 2026-08-28

Updated: 2026-09-22

Comments: 16 pages, 1 figure

License: http://creativecommons.org/licenses/by/4.0/

The gist: Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row- 1 approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued

Terminology

Abstract

Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row- 1 approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued output. Two sharp worst-case laws isolate support geometry: for fixed d and error epsilon, spherical self-attention has rank Θ d, epsilon(n,(1+β)(d-1)/2), while full-ball geometry adds one radial degree and, for β β 0(d, epsilon) and n C d e β/8, gives Θ d, epsilon(β d/2). For a fixed head, row-softmax quotients out row-scalar logit directions: the remaining visible query--key interaction dimension r yields an r/2 per-instance upper law, and bounded constructions show this exponent is minimax sharp. Approximate interaction subspaces incur an explicit residual output error and yield a tolerance-indexed SVD dimension. On an 84-head BERT-base calibration set, we observe modest effective-dimension reductions across many head--temperature settings, together with positive associations with finite constructive rank upper certificates. Together, these results separate support geometry, which sets worst-case temperature scaling, from softmax-visible interaction geometry, which controls per-head approximation complexity.

Sources

Related papers