Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics
Longtian Bao, Jianyou Wang, Yang Zhang, Youze Zheng, Ramamohan Paturi
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable, and models often
Terminology
Abstract
Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable, and models often exhibit an apparent ceiling beyond which additional data yields no further improvement. We study these difficulties in a controlled setting, fine-tuning Qwen2.5-Math-7B on competition mathematics (AIME), a task on which it initially solves only 5.6% of problems (pass@1). To address data scarcity, we introduce Question-begets-Question (QbQ), a scalable procedure in which a teacher transforms existing problems into diverse variants that probe the same underlying skills; to model the absence of oracle reasoning, we train exclusively via reinforcement learning on problem statements and final answers, never on teacher reasoning traces. Static training on such data, however, plateaus well short of the task: real-plus-synthetic augmentation and non-curriculum QbQ generated synthetic data training cap pass@1 at 12.5% and 14.5% respectively, despite large increases in data. Our central finding is that this ceiling is not intrinsic to the model. We propose a self-evolving curriculum that, each round, evaluates the current checkpoint, seeds QbQ from the problems it can mostly get right, and trains on the resulting variants; under an identical data budget, this breaks the ceiling and lifts pass@1 to 16.5% with no sign of saturation after 20 rounds. Counterintuitively, we find that models improve when trained on variants of problems they can mostly get right, and that models trained this way go on to solve harder problems never seen during training.
Sources
- Online Difficulty Filtering for Reasoning Oriented Reinforcement Learning
- Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model
- Nudging the Boundaries of LLM Reasoning
- Cog-DRIFT: Exploration on Adaptively Reformulated Instances Enables Learning from Hard Reasoning Problems
- Self-Evolving Curriculum for LLM Reasoning
- Training Verifiers to Solve Math Word Problems
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
- LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
- Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
- Measuring Mathematical Problem Solving With the MATH Dataset
- LoRA: Low-Rank Adaptation of Large Language Models
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- On the Emergence of Implicit Curriculum in RLVR Learning Dynamics
- ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
- Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
- Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
- Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks