Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective
cs.LG, stat.ML
Submitted: 2025-10-04
Updated: 2026-09-01
Terminology
Sources
- What learning algorithm is in-context learning? Investigations with linear models
- N-gram Language Modeling using Recurrent Neural Network Estimation
- Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
- What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
- How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
- Rethinking Attention with Performers
- What Does BERT Look At? An Analysis of BERT's Attention
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- A Priori Estimates of the Population Risk for Two-layer Neural Networks
- Music Transformer
- Approximation Rate of the Transformer Architecture for Sequence Modeling
- Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
- Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
- Deep Learning via Dynamical Systems: An Approximation Perspective
- ResNet with one-neuron hidden layers is a Universal Approximator
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention
- The Expressive Power of Transformers with Chain of Thought
- Limitations of Normalization in Attention Mechanism
- Scalable-Softmax Is Superior for Attention
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks