The Free-Recipe Limit: Every Recipe Effect Measures Which Premise of an Idealised Learner Broke
cs.AI
Submitted: 2026-08-09
Updated: 2026-08-09
Comments: 42 pages, 9 figures, 17 tables (incl. appendices); code, pre-registration files, and result base to be released with the paper
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fix a corpus and send recipe search to infinity: try every order of the skills, every arrangement from blocked to interleaved, every composition, and keep the best.
Terminology
Abstract
Fix a corpus and send recipe search to infinity: try every order of the skills, every arrangement from blocked to interleaved, every composition, and keep the best. Two quantities decide what that search was worth: the diameter of the reachable set it explores, and the resolution at which anyone can tell two endpoints apart. Where the diameter falls below the resolution, no amount of search converts into a decision, and the signature is not an absence of winners but winners that do not survive re-running. We measure this recipe-search wall with 761 fine-tuning runs on 12 base models (0.5B-14B, three pretraining families) over competition-mathematics skills: base checkpoints, supervised fine-tuning under AdamW, exact-match scoring at k=4. Within one coherent domain at fixed volume the three classical freedoms average 0.010-0.021 against a 0.019 floor, and the largest contrast, 0.0619, clears a three-seed resolution and then reads +0.010 and-0.015 on two reruns. The departure with a systematic answer is coherence: halving one pooled corpus and letting the halves write answers under incompatible but equally correct conventions moves arrangement from capability to allocation between conventions, by two orders of magnitude over a same-convention control, and writing the convention into the input switches the phenomenon off. The switch replicates on a second pretraining family and survives an independent re-execution of its own protocol, with a re-execution spread (0.087) smaller than the resolution a search-selected order cell carries (0.144). Order itself is a transient whose sign crosses zero three times inside a single run. Volume, the one lever nobody calls a recipe, is the one that reliably pays. A public scorecard grades all 26 pre-registered claims: 18 supported, 5 failed, 2 untested, 1 mixed.
Sources
- Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
- What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical Reasoning
- Curriculum Learning for LLM Pretraining: An Analysis of Learning Dynamics
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining
- DoGE: Domain Reweighting with Generalization Estimation
- Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
- The Geometry of Sequential Learning: Lie-Bracket Prediction of Transfer Order
- The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
- Gradient Episodic Memory for Continual Learning
- Don't forget, there is more than forgetting: new metrics for Continual Learning
- Continual Learning for Large Language Models: A Survey
- Learning is Forgetting: LLM Training As Lossy Compression
- Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- Interleaved Multitask Learning with Energy Modulated Learning Progress
- What do Language Models Learn and When? The Implicit Curriculum Hypothesis
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- Optimizer Memory Makes Shuffle Order a First-Order Source of Fine-Tuning Noise
- First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
- Mid-Training of Large Language Models: A Survey
- From Instance Selection to Fixed-Pool Data Recipe Search for Supervised Fine-Tuning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection