How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- High-Dimensional Theory of LoRA Fine-Tuning in a Solvable Attention Model
- Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in Practice
- Taming the Sigmoid Bottleneck: Provably Argmaxable Sparse Multi-Label Classification
- Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
- Kimi K3: Open Frontier Intelligence
- Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
- What You Will Gain By Rounding: Theory and Algorithms for Rounding Rank
- SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
- The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers
- Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks