Understanding In-context Learning of Addition via Activation Subspaces
cs.LG, cs.AI, cs.CL
Submitted: 2025-05-08
Updated: 2026-09-17
Comments: Published as a conference paper at COLM 2026. 10 page main body, 4 page references, 20 page appendix
Code: https://github.com/xyVickyHu/addition-subspaces
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: To perform few-shot learning, language models extract signals from a few input-label pairs, aggregate them into a learned prediction rule, and apply this rule to new inputs.
Terminology
Abstract
To perform few-shot learning, language models extract signals from a few input-label pairs, aggregate them into a learned prediction rule, and apply this rule to new inputs. How is this implemented in the forward pass of modern transformer models? To explore this question, we study a structured family of few-shot learning tasks for which the true prediction rule is to add an integer k to the input. We introduce a novel method that localizes the model's few-shot learning ability to only a few attention heads. This method and the findings generalize to four additional task families spanning arithmetic and semantic tasks. We then perform an in-depth analysis of individual heads via dimensionality reduction and decomposition of the heads' output spaces. For example, in Llama-3-8B-Instruct, we reduce the mechanism underlying these tasks to just three attention heads with six-dimensional subspaces, in which four dimensions track the units digit using trigonometric functions with periods 2, 5, and 10, while two dimensions track magnitude using low-frequency components. To deepen our understanding of this mechanism, we also derive a mathematical identity relating the ''aggregator'' and ''extractor'' subspaces of attention heads, allowing us to track the flow of information from individual examples to a final aggregated concept. Our results demonstrate how tracking low-dimensional subspaces of localized heads throughout a forward pass can provide insight into fine-grained computational structures in language models. Our code is available at https://github.com/xyVickyHu/addition-subspaces.
Sources
- What learning algorithm is in-context learning? Investigations with linear models
- Birth of a Transformer: A Memory Viewpoint
- What you can cram into a single vector: Probing sentence embeddings for linguistic properties
- Induction Heads as an Essential Mechanism for Pattern Matching in In-context Learning
- How do Language Models Bind Entities in Context?
- Interpreting the Second-Order Effects of Neurons in CLIP
- How Do Transformers Learn In-Context Beyond Simple Functions? A Case Study on Learning with Representations
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
- Language Models Use Trigonometry to Do Addition
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
- Progress measures for grokking via mechanistic interpretability
- How Transformers Learn Causal Structure with Gradient Descent
- In-context Learning and Induction Heads
- The mechanistic basis of data dependence and abrupt learning in an in-context classification task
- What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
- Transformers learn in-context by gradient descent
- How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
- Which Attention Heads Matter for In-Context Learning?
- Trained Transformers Learn Linear Models In-Context
- Pre-trained Large Language Models Use Fourier Features to Compute Addition
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks