What can linear attention learn from nonlinear teachers in-context?
stat.ML, cs.LG
Submitted: 2026-10-07
Updated: 2026-10-07
Terminology
Sources
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
- In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners
- Pretrained transformer efficiently learns low-dimensional target functions in-context
- Gated Linear Attention Transformers with Hardware-Efficient Training
- What and How does In-Context Learning Learn? Bayesian Model Averaging, Parameterization, and Generalization
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey