Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation
cs.LG, cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: A language model's representation geometry is not predetermined; it evolves as the model runs.
Terminology
Abstract
A language model's representation geometry is not predetermined; it evolves as the model runs. A faithful account of that geometry must capture that dynamic process, and so cannot be based solely on model-independent statistics such as co-occurrence. Here we introduce a mean-field analysis of attention. The average attention from one token to another defines a kernel that carries representations layer to layer and can be iterated through the network to model how the geometry is transformed. We condition this average two ways. Conditioned on a whole corpus, the kernel predicts the average-case evolution of representation geometry. Conditioned instead on a single context, it predicts the expected geometry for that context. A head's departure from that prediction, its mean-field deviation, isolates the context-specific computation that the mean field misses. Under the corpus-conditional reading, the kernel yields an open-loop model: from the input embeddings and the frozen weights alone, we can iterate the kernel and the model's own MLPs over token representations, never consulting a measured deviation at any layer. The resulting prediction is highly accurate. In early training the model and its corpus mean field are indistinguishable. Replace every attention head with its mean field, and the substitution leaves the loss on real text unchanged. Around the onset of induction, the two diverge, and the gap widens as representations become contextualized. Under the context-conditional reading, deviation from the mean field is a task-agnostic measure of context-specific computation. The residual decomposes additively into unusual attention routing and contextualization of the transported values. Across controlled induction and few-shot settings, greater deviation tracks greater reliance on in-context information.
Sources
- Layer Normalization
- Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B
- Analysis of mean-field models arising from self-attention dynamics in transformer architectures with layer normalization
- A Unified Perspective on the Dynamics of Deep Transformers
- How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- A review of matrix scaling and Sinkhorn's normal form for matrices and positive maps
- Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
- Dynamical Mean-Field Theory of Self-Attention Neural Networks
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks