TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping
cs.CL
Submitted: 2026-04-22
Updated: 2026-08-27
Comments: Accepted at the Third Conference on Language Modeling (COLM 2026)
Code: https://github.com/ZubinGou/math-evaluation-harness
License: http://creativecommons.org/licenses/by/4.0/
The gist: The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling LRMs to reason longer, and more accurately.
Terminology
Abstract
The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling LRMs to reason longer, and more accurately. However, a growing body of studies show that LRMs are still inefficient, over-generating verification and reflection steps. Additionally, the high-level role of each reasoning step and how different step types contribute to the generation of correct answers, is largely underexplored. To address this challenge, we introduce TRACES (Tagging of the Reasoning steps enabling Adaptive Cost-Efficient early-Stopping), a lightweight framework that tags reasoning steps in real-time, and enable adaptive, cost-efficient early stopping of large-language-model inferences. By monitoring reasoning behaviors during inferences, we find that LRMs tend to shift their reasoning behavior after reaching a correct answer. We demonstrate that the monitoring of the specific type of steps can produce effective interpretable early stopping criteria. Evaluated on three mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and two knowledge and reasoning benchmarks (MMLU and GPQA), TRACES achieve 20 to 50% token reduction while maintaining comparable accuracy to standard generation. On harder tasks (BeyondAIME, IMO AnswerBench), the accuracy-efficiency trade-off is steeper, requiring a more conservative early-stopping threshold. This work offers a novel way to study and control the generation behavior of LRMs.
Sources
- Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
- Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
- Evaluating Large Language Models Trained on Code
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Training Verifiers to Solve Math Word Problems
- Complexity-Based Prompting for Multi-Step Reasoning
- I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
- Token-Budget-Aware LLM Reasoning
- Measuring Mathematical Problem Solving With the MATH Dataset
- From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
- What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
- How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
- Evaluating Step-by-step Reasoning Traces: A Survey
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- Let's Verify Step by Step
- AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
- Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
- DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering