Post-Reasoning: Improving the Performance of Non-Thinking Models at No Cost
cs.AI
Submitted: 2026-05-07
Updated: 2026-09-14
Code: https://github.com/project-numina/aimo-progress-prize
License: http://creativecommons.org/licenses/by/4.0/
The gist: As the widespread adoption of Large Language Models (LLMs) accelerates, token consumption from intermediate reasoning traces increasingly contributes to inference latency and operational cost.
Terminology
Abstract
As the widespread adoption of Large Language Models (LLMs) accelerates, token consumption from intermediate reasoning traces increasingly contributes to inference latency and operational cost. Recent studies suggest that many real-world tasks require little to no explicit reasoning, with additional reasoning sometimes even degrading performance. In this work, we propose Post-Reasoning, a simple yet effective approach that improves instruction-tuned models by conditioning them to justify their answers after generating the final response. By design, it enables the final answer to be obtained without additional latency or token cost, while still improving performance through simple instruction augmentation. We evaluate Post-Reasoning across 117 model--benchmark settings spanning 13 open and proprietary models, 4 model families, and 9 diverse reasoning and knowledge-intensive benchmarks, including AMC, HMMT, GSM8K, GPQA, MMLU-Pro, and BIG-Bench Hard. Post-Reasoning improves performance in over 88.19% of evaluated settings, achieving a mean relative improvements of 17.37%. Furthermore, we propose supervised post-reason tuning, which further improves performance in over 91.11% of evaluated settings, and exceeds the prompt-based post-reasoning baseline by an average of 8.01%, demonstrating that post-reasoning can be effectively internalized through training. Ultimately, Post-Reasoning establishes a new performance ceiling for direct-answer capabilities.
Sources
- OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
- The Llama 3 Herd of Models
- Token-Budget-Aware LLM Reasoning
- Training Large Language Models to Reason in a Continuous Latent Space
- LoRA: Low-Rank Adaptation of Large Language Models
- Datasets: A Community Library for Natural Language Processing
- ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy
- Ministral 3
- SGDR: Stochastic Gradient Descent with Warm Restarts
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
- CoT-Valve: Length-Compressible Chain-of-Thought Tuning
- Orca-Math: Unlocking the potential of SLMs in Grade School Math
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection