Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
Yuto Nishida, Hirokazu Kiyomaru, Yusuke Oda, Takashi Kodama, Chaoran Liu, Daisuke Kawahara, Yusuke Miyao, Max Müller-Eberstein, Masaru Isonuma
Nara Institute of Science and Technology · NII LLMC · Waseda University · The University of Tokyo · IT University of Copenhagen · Tohoku University
cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Accepted to COLM 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper proposes a task-agnostic measure of training data influence for language model pretraining.
Terminology
Summary
This paper proposes a task-agnostic measure of training data influence for language model pretraining. The key innovation is reformulating training data influence without requiring a downstream task or validation set as the attribution target. Instead, the authors define an example's influence by how much its gradient update reduces the squared distance to the final parameters of a given pretraining run.
The method defines influence at two levels:
Mini-batch influence: For updates indexed by t, with parameters θt and final parameters θ*, the influence of mini-batch Bt is defined as the reduction in squared L2 distance to the final parameters: Cont(Bt) = St − St+1, where St = ∥θ* − θt∥22. This expands to Cont(Bt) = 2∆⊤t(θ* − θt) − ∥∆t∥22, where the first term measures alignment between the update direction and the direction toward final parameters, and the second term penalizes the squared norm of the update.
Example-level influence: Under an SGD assumption, the influence of example xkt is defined as Cont(xkt):= 2∆⊤t,k(θ* − θt) − ∆⊤t,k∆t, which sums to the mini-batch influence. The authors note a connection to TracIn (Pruthi et al., 2020) under the simplifying assumption of pairwise orthogonal per-example update vectors.
Checkpoint-based approximation: Since exact computation requires access to all intermediate states, the authors approximate using periodically saved checkpoints: Cont(xkt) ≈ 2∆⊤c,k(θ* − θc) − ∆⊤c,k(θc′ − θc), where c and c′ are consecutive checkpoint steps.
The method is applied to:
-
Pythia suite (Biderman et al., 2023): six scales (70M, 160M, 410M, 1.4B, 6.9B, 12B), each trained on 300B tokens from The Pile with 154 checkpoints
-
PolyPythias (van der Wal et al., 2025): 160M and 410M scales with variations in model initialization and data ordering
Each 2048-token sequence is treated as one training example. Examples are sampled from the actual training data stream associated with each checkpoint interval.
The mean contribution increases substantially after the early stage, becomes largest around roughly 40k steps, and then gradually decreases toward the end of training.
The standard deviation follows a similar pattern. The share of opponent
examples (those with negative contribution, moving the model away from final parameters) "remains relatively rare through the early and middle stages, staying around 0–1% for most of the trajectory before about 70k steps. After that point, however, their share rises markedly, reaching roughly 6% around 90k steps and around 8–10% near the end of training."
Using perplexity (PPL) computed with the final checkpoint of Pythia-12B-Deduped, the authors find that higher-PPL examples account for a larger share of normalized contribution during the middle stage of pretraining, while contribution is more evenly distributed across PPL bins during the early and late stages.
Specifically, the share of the highest-PPL bin increases from approximately 20% to 25% in the middle stage, whereas that of the lowest-PPL bin decreases from approximately 20% to 15%.
The most striking finding is a systematic temporal shift in which domains contribute most:
-
Early training:
literature-related data are more strongly aligned with the trajectory toward the final parameters
-
Later stages:
STEM data become more strongly aligned in later stages
Specifically, for the bottom 5% of contributors, STEM-related domains such as Computers and Electronics and Science account for a large proportion
during the first half of training, but in the later stages, however, the shares of low-contribution STEM domains decrease, while domains such as Books and Literature become more common among the low contributors.
For the top 5%, STEM domains are not necessarily dominant among high-contribution examples early in training, but their share tends to increase in later stages.
The authors validated the checkpoint-based approximation by retraining selected intervals of Pythia-70M-Deduped with all intermediate states stored. Correlations between exact and approximate contributions were: early interval (1k→2k): Pearson r = 0.622, Spearman r = 0.592; middle interval (10k→11k): Pearson r = 0.808, Spearman r = 0.811; late interval (100k→101k): Pearson r = 0.950, Spearman r = 0.947. The approximation becomes substantially stronger later in training
because parameter changes between adjacent checkpoints are larger early in training and become smaller later.
Using alternative reference checkpoints at steps 30k, 70k, 120k, and 140k (compared to the final 143k checkpoint), the authors found: "Using the near-final 140k checkpoint as the reference leaves the contribution rankings almost unchanged, yielding a mean Spearman correlation of 0.997 and a mean top-5% overlap of 0.972 with the original scores. By contrast, substantially earlier reference checkpoints produce markedly different rankings."
-
Model sizes:
Across the six Pythia-Deduped model sizes, the contribution-distribution, text-difficulty, and domain analyses exhibit broadly similar stage-dependent dynamics, while also showing systematic scale-dependent differences.
Notably,several transitions that occur during the middle stage for larger models appear later for smaller models.
-
Weight initialization and data ordering:
The qualitative domain-level contribution dynamics are also broadly consistent across PolyPythia runs with different random factors,
with the exception of an anomalous 410M seed-3 run identified as an outlier with loss spikes.Changing data ordering introduces somewhat greater variation
than weight initialization.
The authors compared their measure with a TracIn-style influence score computed with respect to domain-wise language-modeling loss on held-out validation sets. They found that the resulting attribution patterns depend substantially on the selected validation domain
and that the domain-specific TracIn analysis does not consistently recover the Literature-to-STEM crossover observed with our contribution measure.
The authors identify that higher-PPL texts and STEM-related domains tend to contribute less during the initial learning phase but more strongly during the critical learning phase.
They note that their results lend empirical support
to recent LLM pretraining recipes that increase the allocation of STEM-related data, such as math and code, in later stages of training.
They also argue that their findings motivate pretraining strategies that adapt data composition across training stages rather than treating a single data mixture as equally suitable throughout training.
The authors acknowledge several limitations: the measure is defined relative to the final parameters of a particular pretraining run
; squared parameter-space distance is sensitive to parameterization and does not directly measure changes in model behavior
; the checkpoint approximation is only validated on three intervals of Pythia-70M-Deduped; the example-level decomposition assumes SGD while actual training uses adaptive optimizers; all experiments are within closely related model families and pretraining data
; and the analyses characterize how contribution varies across training rather than directly establishing the causal effect of changing the data mixture.
Improvements for AI systems
Improvement 1: Dynamic Data-Curriculum Scheduling in Pretraining
- The improved AI system can automatically adjust the mixture of training data domains over time based on the measured contribution trajectory. Instead of using a static data mixture, the system will increase the sampling weight of STEM-related data (math, code, science) during the middle-to-late stages of training, while front-loading literature and general text in early stages. This is directly informed by the observed Literature-to-STEM crossover. The system can implement a scheduler that queries the contribution measure at each checkpoint interval and re-weights the next batch’s domain composition to maximize alignment with the final parameter trajectory.
Improvement 2: Early-Stopping and Checkpoint Selection via Contribution Trajectory
- The improved AI system can use the contribution distribution (mean, standard deviation, and opponent-example share) as a real-time training health monitor. When the share of negative-contribution examples exceeds a threshold (e.g., >8%) or when mean contribution begins to decline after the peak (40k steps in Pythia), the system can trigger early stopping or switch to a fine-tuning phase. This prevents wasted compute on late-stage updates that move parameters away from the optimal final state. The system can also select the best checkpoint for downstream deployment by identifying the step where mean contribution is maximal, rather than using the final checkpoint.
Improvement 3: Adaptive Learning-Rate and Optimizer Adjustment Based on Per-Example Influence
- The improved AI system can use the example-level contribution score (Cont(x kt)) to modulate the effective learning rate for individual examples in real time. Examples with negative contribution (opponents) can be down-weighted or excluded from updates, while high-contribution examples (especially those aligned with the final parameter direction) can be up-weighted. This can be implemented as a per-example gradient scaling factor: multiply the gradient of example x by a sigmoid of its contribution score. This reduces the 8–10% of late-training opponent examples that currently push the model away from its final parameters, improving final model quality without additional data.
Improvement 4: Domain-Specific Data Augmentation and Synthetic Data Generation
- The improved AI system can identify which domains are under-contributing in the early stage (e.g., STEM) and generate synthetic training examples in those domains to accelerate their contribution alignment. Using the contribution measure as a reward signal, the system can train a small generative model to produce new STEM-like texts that maximize the predicted contribution score when added to the training stream. This is particularly useful for the early-to-middle transition where STEM data’s contribution is low but rising—synthetic augmentation can steepen that rise, shortening the critical learning phase.
Improvement 5: Robustness to Random Initialization and Data Ordering
- The improved AI system can use the contribution measure as a diagnostic to detect anomalous training runs (like the 410M seed-3 outlier with loss spikes). By monitoring the domain-level contribution dynamics across multiple seeds, the system can flag runs where the expected Literature-to-STEM crossover does not occur or where opponent-example share spikes abnormally early. This enables automatic retraining with a different seed or data ordering, improving the reliability of pretraining pipelines. The system can also use the contribution measure to select the most consistent seed/ordering combination before full-scale training, by running short 1k-step probes and comparing their contribution trajectories.
Improvement 6: Task-Agnostic Data Filtering for Downstream Fine-Tuning
- The improved AI system can use the contribution measure (relative to a pretraining run’s final parameters) as a filter for selecting which pretraining examples to carry into a fine-tuning phase. Examples with high contribution in the late stage (e.g., STEM, high-PPL) are likely to be more useful for transfer to reasoning-heavy downstream tasks. The system can rank all pretraining examples by their contribution score and create a prioritized subset for continued training or for building task-specific adapters, reducing fine-tuning data requirements while maintaining or improving performance. This is more principled than random sampling or perplexity-based filtering.
Abstract
Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training. We propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target. Specifically, we define an example's influence by how much its gradient update reduces the squared distance to the final parameters of a given pretraining run, and estimate this quantity from intermediate checkpoints without retraining. Applying the method to 18 configurations from the Pythia and PolyPythia suites, we find systematic temporal changes in influential data. Early in training, literature-related data are more strongly aligned with the trajectory toward the final parameters, whereas STEM data become more strongly aligned in later stages. This qualitative crossover is broadly consistent across model configurations. Our results provide a tractable trajectory-level view of how influential data change throughout pretraining, complementing influence analyses defined with respect to specific downstream tasks or validation sets.
Sources
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering