When2Think: Learning When and How Much to Reason
cs.AI, cs.LG
Submitted: 2026-09-17
Updated: 2026-09-28
Code: https://github.com/huggingface/math-verify
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones.
Terminology
Abstract
Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on easy instances for accuracy loss on hard instances. We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism that leverages pre-computed reference statistics (accuracy and token usage) to regulate reasoning depth. Combined with verifier-based rewards and batch-wise standardized advantages, IDAC enables stable critic-free optimization without learned reward models or online reference-model queries. When2Think encourages direct answering on easy instances while preserving extended reasoning on hard instances, thereby learning when to use System 1 (NoThink) versus System 2 (Think). Experiments on mathematical benchmarks demonstrate improved accuracy-efficiency trade-offs: on AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model, and on AIME25, When2Think achieves 40.0% Pass@3, outperforming compression and routing-only baselines.
Sources
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- A Survey on Post-training of Large Language Models
- LLM Post-Training: A Deep Dive into Reasoning Large Language Models
- Optimizing Length Compression in Large Reasoning Models
- ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
- OpenAI o1 System Card
- Phi-4-reasoning Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Reasoning Language Models: A Blueprint
- A Survey of Reinforcement Learning for Large Reasoning Models
- Training Verifiers to Solve Math Word Problems
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
- Evaluating Large Language Models Trained on Code
- Qwen2.5 Technical Report
- Olmo 3
- Distilling the Knowledge in a Neural Network
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection