Understanding the Staged Dynamics of Transformers in Learning Latent Structure
cs.LG
Submitted: 2025-11-24
Updated: 2026-09-16
Comments: Preprint
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language modeling has shown us that transformers can discover latent structure from context, but the dynamics of how they acquire different components of that structure remain poorly understood,
Terminology
Abstract
Language modeling has shown us that transformers can discover latent structure from context, but the dynamics of how they acquire different components of that structure remain poorly understood, leading to assertions that models just remix training data. In this work, we use the Alchemy benchmark in a controlled setting (Wang et al.,2021) to investigate latent structure learning. We train a small decoder-only transformer on three task variants: 1) inferring missing transitions from partial contextual information, 2) composing simple rules to solve multi-transition sequences, and 3) decomposing complex multi-step examples to infer intermediate transitions. By factorizing each task into interpretable components, we show that the model learns the different latent structure components in discrete stages. We also observe an asymmetry: the model composes fundamental transitions robustly, but struggles to decompose complex examples to discover the atomic transitions. Finally, using causal interventions, we identify layer-specific plasticity windows during which freezing substantially delays or prevents stage completion. These findings provide insight into how a transformer model acquires latent structure, offering a detailed view of how capabilities evolve during training.
Sources
- Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- When can transformers compositionally generalize in-context?
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- SGD on Neural Networks Learns Functions of Increasing Complexity
- How Transformers Learn Causal Structure with Gradient Descent
- In-context Learning and Induction Heads
- Progress measures for grokking via mechanistic interpretability
- The mechanistic basis of data dependence and abrupt learning in an in-context classification task
- Emergent Abilities of Large Language Models
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Training Dynamics of In-Context Learning in Linear Attention
- Alchemy: A benchmark and analysis toolkit for meta-reinforcement learning agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks