Improving Mathematical Reasoning Capabilities in Large Language Models via Reasoning Process Error Classification
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 10 pages
Code: https://github.com/deepseek-ai/DeepSeekMath
License: http://creativecommons.org/licenses/by/4.0/
The gist: The reasoning ability of large language models (LLMs) is a critical factor for practical LLM-based applications.
Terminology
Abstract
The reasoning ability of large language models (LLMs) is a critical factor for practical LLM-based applications. To investigate the current reasoning capability of LLMs, we clarify the types of errors that arise in LLMs' reasoning processes on mathematical datasets. We focus on problems where LLMs produce an incorrect answer. We define errors in the reasoning process as reasoning errors and manually analyze the features of reasoning errors. We defined and classified 21 error classes and identified the frequently occurring classes among them. Beyond qualitative evaluation, we leverage the evaluation results to improve the reasoning capability. We designed a prompt that explicitly focuses on eight error classes. The experiments demonstrate that this prompt effectively improves reasoning performance. Furthermore, the results suggest that the frequent reasoning errors identified in this paper are common across LLMs of comparable scale.
Sources
- Large Language Models and Mathematical Reasoning Failures
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Training Verifiers to Solve Math Word Problems
- Gemma 2: Improving Open Language Models at a Practical Size
- GPT-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering