Test time training enhances in-context learning of nonlinear functions
stat.ML, cs.LG
Submitted: 2025-09-30
Updated: 2026-09-10
Comments: Under review at NeurIPS 2026. 44 pages, 2 figures, appendix included; revised synthetic experiment, corrected mistakes, and added background section
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- Learning Single-Index Models with Shallow Neural Networks
- The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Learning Capabilities
- Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions
- Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
- Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
- Neural Networks can Learn Representations with Gradient Descent
- Linearized two-layers neural networks in high dimension
- Computational-Statistical Gaps in Gaussian Single-Index Models
- How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
- In-Context Convergence of Transformers
- Learning without training: The implicit dynamics of in-context learning
- Many-Shot In-Context Learning in Multimodal Foundation Models
- On the Complexity of Learning Sparse Functions with Statistical and Gradient Queries
- Adam: A Method for Stochastic Optimization
- Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
- What Can Transformers Learn In-Context? A Case Study of Simple Function Classes
- In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
- Test-Time Training Done Right
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey