Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum
cs.LG, stat.ML
Submitted: 2026-03-18
Updated: 2026-08-27
Comments: 39 pages, 4 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chain-of-thought reasoning, where language models expend additional computation by producing thinking tokens prior to final responses, has driven significant advances in model capabilities.
Terminology
Abstract
Chain-of-thought reasoning, where language models expend additional computation by producing thinking tokens prior to final responses, has driven significant advances in model capabilities. However, training these reasoning models is extremely costly in terms of both data and compute, as it involves collecting long traces of reasoning behavior from humans or synthetic generators and further post-training the model via reinforcement learning. Are these costs fundamental, or can they be reduced through better algorithmic design? We show that autocurriculum, where the model uses its own performance to decide which problems to focus training on, provably improves upon standard training recipes for both supervised fine-tuning (SFT) and reinforcement learning (RL). For SFT, we show that autocurriculum requires exponentially fewer reasoning demonstrations than non-adaptive fine-tuning, by focusing teacher supervision on prompts where the current model struggles. For RL fine-tuning, autocurriculum decouples the computational cost from the quality of the reference model, reducing the latter to a burn-in cost that is nearly independent of the target accuracy. These improvements arise purely from adaptive data selection, drawing on classical techniques from boosting and learning from counterexamples, and requiring no assumption on the distribution or difficulty of prompts.
Sources
- Phi-4-reasoning Technical Report
- The Coverage Principle: How Pre-Training Enables Post-Training
- Evaluating Large Language Models Trained on Code
- Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
- Prompt Curriculum Learning for Efficient LLM Post-Training
- On the Statistical Query Complexity of Learning Semiautomata: a Random Walk Approach
- OpenThoughts: Data Recipes for Reasoning Models
- Reinforced Self-Training (ReST) for Language Modeling
- On The Power of Curriculum Learning in Training Deep Networks
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- OpenAI o1 System Card
- A Theory of Learning with Autoregressive Chain of Thought
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
- Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- What Objective Does Self-paced Learning Indeed Optimize?
- h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
- Olmo 3
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks