Gaussian Equivalence for Multi-Head Self-Attention
stat.ML, cs.LG, math.PR
Submitted: 2026-10-07
Updated: 2026-10-07
Terminology
Sources
- The Falcon Series of Open Language Models
- Language Models are Few-Shot Learners
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Effective Theory of Transformers at Initialization
- On the Interpolation Error of Nonlinear Attention versus Linear Regression
- Fast Transformer Decoding: One Write-Head is All You Need
- DINOv3
- Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey