Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA

arXiv:2605.07959 · cs.LG, math.FA, math.PR · Submitted 2026-05-08 · Read on arXiv

cs.LG, math.FA, math.PR

Submitted: 2026-05-08

Updated: 2026-09-06

License: http://creativecommons.org/licenses/by/4.0/

The gist: Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications.

Terminology

Abstract

Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby a factorized parameterization of them is trained, to achieve a surprisingly beneficial accuracy-size trade-off. In this work, via a unified framework we rigorously establish trainability of such models under stochastic methods. We prove that for a class of mild regularizations, the empirical regression loss on a attention layer and LoRA on a shallow neural net, both induce Poincaré inequality for the corresponding Gibbs' measure. Crucially, we show that the Poincaré constant is free of the data dimension for LoRA and for multi-head attention it is free of the head dimensions, Then it follows, via invoking recent results, that a certain SDE, which mimics the SGD, minimizes the corresponding losses. In both the cases, our first-of-its-kind results of trainability on attention and nets, do not use any assumptions on the data or the model size.

Sources

Related papers