Optimizer Memory Makes Shuffle Order a First-Order Source of Fine-Tuning Noise
John Sweeney
cs.LG, math.NA, math.OC, stat.ML
Submitted: 2026-06-28
Comments: 29 pages, 3 figures, 12 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Shuffle order can be a larger source of fine-tuning noise than a memoryless analysis predicts: fixed-clock optimizer memory makes local equal-multiset contrasts first order in the learning rate
Terminology
Abstract
Shuffle order can be a larger source of fine-tuning noise than a memoryless analysis predicts: fixed-clock optimizer memory makes local equal-multiset contrasts first order in the learning rate rather than second order, and the resulting order channel can be large enough for a single seed to flip a close A/B comparison. We isolate this mechanism and derive a fit-free way to size the noise it produces. For a memoryless optimizer, reordering an equal multiset has no first-order endpoint term; the leading local contrast is the O(eta 2) gradient bracket. Fixed-clock optimizers such as AdamW are different. Their moment buffers, preconditioner state, and de-biasing counters advance with the step index rather than with the learning-rate-scaled time tau= eta k, so the same gradient can receive a position-dependent endpoint weight. For any fixed finite measurement window, a lifted-state expansion gives an O(eta) equal-multiset contrast whenever the first-order replay coefficient is nonzero, while regular and clock-matched controls remain O(eta 2); a bare fixed- beta momentum buffer is already enough. A bitwise-deterministic replay from one warmed optimizer state isolates the mechanism, giving order-variance slopes 1.83 for AdamW, 2.00 for fixed- beta momentum, and 4.00 for SGD; matching the memory clock to tau restores the regular exponent. For AdamW with a frozen preconditioner, the same impulse-weight kernel gives a closed-form asymptotic order-variance floor after the local potentials are measured, with no fitted coefficients. The result is local to the measurement window (independent reshuffling can average the channel across windows), but it yields order-noise error bars, positional attribution weights, and a seed-budget criterion for fine-tuning comparisons.
Sources
- SGD with shuffling: optimal rates without component convexity and large epoch requirements
- On the Trajectories of SGD Without Replacement
- How Memory in Optimization Algorithms Implicitly Modifies the Loss
- Does the Order of Fine-tuning Matter and Why?
- Symbolic Discovery of Optimization Algorithms
- Implicit biases in multitask and continual learning from a backward error analysis perspective
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- Computing the Variance of Shuffling Stochastic Gradient Algorithms via Power Spectral Density Analysis
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- Why Random Reshuffling Beats Stochastic Gradient Descent
- Scaling Laws for Neural Language Models
- Fresh in memory: Training-order recency is linearly encoded in language model activations
- The Order Is The Message
- GraB: Finding Provably Better Data Permutations than Random Reshuffling
- On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
- Closing the convergence gap of SGD without replacement
- Commute Your Domains: Trajectory Optimality Criterion for Multi-Domain Learning
- Process-Tensor Tomography of SGD: Measuring Non-Markovian Memory via Back-Flow of Distinguishability
- Without-Replacement Sampling for Stochastic Gradient Methods: Convergence Results and Application to Distributed Optimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks