Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road
cs.LG
Submitted: 2026-05-16
Updated: 2026-09-01
Comments: 26 pages, 16 figures
Code: https://github.com/psunlpgroup/reasoning_forks
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through specialized fine-tuning procedures.
Terminology
Abstract
Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through specialized fine-tuning procedures. While these methods reliably improve pass@1 accuracy, prior works have observed that they show a coverage shrinkage behavior, where pass@k degrades relative to the base model. In this paper, we investigate the cause of reasoning shrinkage under SFT-based post-training. We hypothesize that this behavior is driven by properties of the fine-tuning data, specifically related to decision points or "forks in the road" scenarios where model encounters indecipherable patterns with multiple valid reasoning paths. To test this hypothesis, we design controlled case studies that simulate such decision-point settings, spanning indecipherable nodes in graph branching, and reasoning modes. By tracking post-training dynamics in these settings, we find that the shrinkage phenomenon is tightly correlated with the prevalence of decision-point scenarios in the training data. We also demonstrate that this shrinkage behavior can be partially mitigated through targeted data synthesis design of decision-points and a more systematic diversity-encouraging decoding mechanism. Our findings identify data-centric factors as a key driver of shrinkage in reasoning models and highlight diversity-aware designs as an effective lever for controlling it. (Data and code for reproducing our experiments are available at https://github.com/psunlpgroup/reasoning forks)
Sources
- Training Verifiers to Solve Math Word Problems
- Skill-Targeted Adaptive Training
- Mixtral of Experts
- Let's Verify Step by Step
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- EvoLM: In Search of Lost Language Model Training Dynamics
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
- Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
- Unveiling Transformers with LEGO: a synthetic reasoning task
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks