Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA
cs.LG, math.FA, math.PR
Submitted: 2026-05-08
Updated: 2026-09-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications.
Terminology
Abstract
Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby a factorized parameterization of them is trained, to achieve a surprisingly beneficial accuracy-size trade-off. In this work, via a unified framework we rigorously establish trainability of such models under stochastic methods. We prove that for a class of mild regularizations, the empirical regression loss on a attention layer and LoRA on a shallow neural net, both induce Poincaré inequality for the corresponding Gibbs' measure. Crucially, we show that the Poincaré constant is free of the data dimension for LoRA and for multi-head attention it is free of the head dimensions, Then it follows, via invoking recent results, that a certain SDE, which mimics the SGD, minimizes the corresponding losses. In both the cases, our first-of-its-kind results of trainability on attention and nets, do not use any assumptions on the data or the model size.
Sources
- Longformer: The Long-Document Transformer
- Generating Long Sequences with Sparse Transformers
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Adam: A Method for Stochastic Optimization
- LoRA Training Provably Converges to a Low-Rank Global Minimum or It Fails Loudly (But it Probably Won't Fail)
- Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
- Generative AI for fast and accurate statistical computation of fluids
- What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
- FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators
- Linformer: Self-Attention with Linear Complexity
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks