Prescriptive SVD-Inspired Attention via Spectral Energy Retention

arXiv:2609.24370 · cs.LG, cs.CV · Submitted 2026-09-21 · Read on arXiv

cs.LG, cs.CV

Submitted: 2026-09-21

Updated: 2026-09-21

Comments: Published in Transactions on Machine Learning Research (TMLR), 2026

Journal ref: Transactions on Machine Learning Research, 2026

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Self-attention is central to modern Transformer architectures, but its dense dot-product formulation makes it difficult to identify which internal directions are structurally important and which can

Terminology

Abstract

Self-attention is central to modern Transformer architectures, but its dense dot-product formulation makes it difficult to identify which internal directions are structurally important and which can be modified without disrupting the model. SVD-Inspired Attention (SVDA) addresses part of this problem by introducing a learned diagonal spectrum into the query-key score interaction, making latent attention directions explicitly inspectable through indicators such as spectral entropy, effective rank, sparsity, alignment, selectivity, and perturbation response. This paper examines the transition from diagnostic interpretation to operational intervention. A diagnosis--intervention--verification framework is proposed, and one intervention is evaluated: spectral energy retention in the attention-score pathway. Across FashionMNIST, CIFAR-10, CIFAR-100, and Food-101, the ρ=0.90 prescription removes 24.5--53.7% of score directions, reduces parameters by 2.6--4.3%, and reduces estimated MACs by 2.8--5.4%. The paired mean accuracy change of the dimension-reduced model ranges from-0.03 to +0.05 percentage points over three seeds. These results support SVDA as an intrinsically interpretable attention mechanism whose learned spectrum exposes an operational coordinate system for deterministic and verifiable modification of attention-score formation.

Sources

Related papers