Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning
stat.ML, cs.LG
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- In-Context Operator Learning on the Space of Probability Measures
- Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
- Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer
- Transformers Meet In-Context Learning: A Universal Approximation Theory
- Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization
- Transformers for Learning on Noisy and Task-Level Manifolds: Approximation and Generalization Insights
- Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
- Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey