Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training
cs.LG
Submitted: 2026-08-19
Updated: 2026-09-07
Comments: v2: abstract metadata formatting fix only; paper unchanged
Code: https://github.com/zacharyspeck/burst-study
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: A single training example's contribution to a finished model is normally estimated rather than measured, because measuring it takes two expensive full pre-training runs that differ in one row of one
Terminology
Abstract
A single training example's contribution to a finished model is normally estimated rather than measured, because measuring it takes two expensive full pre-training runs that differ in one row of one batch. We ran that counterfactual 24 times at a small scale. We trained 32 GPT-2 models at 124M parameters from scratch on OpenWebText, over four conditions and eight seeds. At step 200 of 9,536, at peak learning rate, we replaced one row of a 256-row batch with a fixed context injection carrying a 194-token passage. The three injected conditions are: 1. fluent prose with a corpus-attested subject, 2. fluent prose with a fabricated subject matched to it within 0.14% on full-batch gradient delta, and 3. random keyboard characters. The fourth condition is an uninjected twin. The passage is learned from one exposure and then decays. Fifty steps after injection, the arm that saw a passage predicts it better than the arm that did not by 0.039 and 0.044 nats of cross-entropy on the passage, at eight of eight seeds with p < 0.0001. At the final step we do not detect that difference for either passage, at p = 0.25 and p = 0.71, against minimum detectable effects of 0.025 and 0.079 nats, nor between the two passages, at p = 0.54. Every geometric measure we report is taken after that decay. Our pre-registered contrast on interpolation loss barrier is +0.0068 with p = 0.509, against a minimum detectable effect of 0.032 barrier units. Held-out cross-entropy is-0.00044 with p = 0.310. Per-layer centered kernel alignment does not detectably separate any condition at any layer. Weight displacement reaches 44.1% of the seed-to-seed Euclidean distance and is 92% settled by the midpoint of training, while the barrier reaches 3.0% of the seed-to-seed barrier. Those two figures sit roughly 15 times apart, and that is a lower bound. The injection relocates the model within its basin without moving it out.
Sources
- Critical Learning Periods in Deep Neural Networks
- Git Re-Basin: Merging Models modulo Permutation Symmetries
- If Influence Functions are the Answer, Then What is the Question?
- Accounting for Variance in Machine Learning Benchmarks
- How Do Large Language Models Acquire Factual Knowledge During Pretraining?
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks
- Linear Mode Connectivity and the Lottery Ticket Hypothesis
- Studying Large Language Model Generalization with Influence Functions
- Linear Mode Connectivity under Data Shifts for Deep Ensembles of Image Classifiers
- Datamodels: Predicting Predictions from Training Data
- Understanding Black-box Predictions via Influence Functions
- The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
- Towards falsifiable interpretability research
- TRAK: Attributing Model Behavior at Scale
- Estimating Training Data Influence by Tracing Gradient Descent
- PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks