It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-03
Updated: 2026-09-03
Comments: 25 pages, 11 figures, 4 tables. Also available at doi:10.5281/zenodo.22261107. Code and pre-registered protocols: https://github.com/bulutyigit/problem-not-path
Code: https://github.com/bulutyigit/problem-not-path
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates.
Terminology
Abstract
Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the continuation budget from when a prefix carries value that fresh computation cannot buy, comparing per-anchor continuation solve rates against from-scratch restart curves at matched total generated-token budget. Applied to 178 problem-model cells (89 MATH problems x two small open models, an outcome-blind but difficulty-targeted cohort), exactly 1 of 178 cells survives as prefix-limited; restart dose-response separates a compute-starved model from a capability-limited one; and wherever the matched budget lies inside the restart grid, continuing the model's own prefix beats restarting (9 of 9) -- predominantly compute compression rather than expanded reachability. Second, a pre-registered, difficulty-controlled test finds no detectable outcome information in early-window internal signals beyond a problem-difficulty baseline, and two generation-free analyses of public corpora show why this control is needed: a trace-blind difficulty proxy reaches AUROC 0.873 on 192K DeepSeek-R1 generations -- inside the published probe range -- and a closely matched reconstruction of the closest published early-window positive recovers a comparable pooled result (0.849) while within problem it is statistically indistinguishable from chance at all ten anchors (0.496 at t=4); a post-hoc within-targeted probe finds only a small average residual, concentrated in three low-failure problems. High pooled probe AUROCs cannot by themselves establish within-attempt information; a question-only baseline or within-problem evaluation is required.
Sources
- Phi-4 Technical Report
- Probing the Trajectories of Reasoning Traces in Large Language Models
- Forking Paths in Neural Text Generation
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- The Illusion of Insight in Reasoning Models
- Temporal Predictors of Outcome in Reasoning Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Efficiently Scaling LLM Reasoning with Certaindex
- Deep Think with Confidence
- Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning
- Inverse Scaling in Test-Time Compute
- Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
- Reasoning Models Don't Just Think Longer, They Move Differently
- Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them)
- First Try Matters: Revisiting the Role of Reflection in Reasoning Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks