Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces
cs.AI, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD).
Terminology
Abstract
Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a self-evolving curriculum. LoT automatically generates semantically faithful but easier variants of reasoning problems, organizes them into difficulty buckets using step-based measures, and employs a self-evolving bandit scheduler to allocate training adaptively. Evaluated on two reasoning domains, math and multi-hop reasoning, across 1-8B models from different families, LoT consistently improves over KD. It delivers large gains on arithmetic tasks (e.g., +32 percentage points on AddSub, +25pp on SVAMP), +2-8pp improvements on in-domain test splits, and strong though dataset-dependent benefits on multi-hop reasoning (e.g., +16pp on QASC, +25pp on StrategyQA). LoT also converges faster than staged curricula, highlighting the value of adaptive progression. These results show that progressive rewrites coupled with adaptive curricula provide a simple yet effective recipe for strengthening reasoning in smaller LLMs.
Sources
- Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational Agents
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Self-Evolving Curriculum for LLM Reasoning
- Training Verifiers to Solve Math Word Problems
- Explaining Answers with Entailment Trees
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- MiniLLM: On-Policy Distillation of Large Language Models
- Distilling the Knowledge in a Neural Network
- Large Language Models Are Reasoning Teachers
- The Impact of Reasoning Step Length on Large Language Models
- QASC: A Dataset for Question Answering via Sentence Composition
- Small Models Struggle to Learn from Strong Reasoners
- TinyGSM: achieving >80% on GSM8k with small language models
- Let's Learn Step by Step: Enhancing In-Context Learning Ability with Curriculum Learning
- Teaching Small Language Models to Reason
- Orca 2: Teaching Small Language Models How to Reason
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
- Are NLP Models really able to Solve Simple Math Word Problems?
- Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
- LADDER: Self-Improving LLMs Through Recursive Problem Decomposition
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection