State of Thought Enables Endogenous Reasoning
cs.CL, cs.AI
Submitted: 2026-09-13
Updated: 2026-09-25
License: http://creativecommons.org/licenses/by/4.0/
The gist: Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs).
Terminology
Abstract
Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through costly expansion in constrained search spaces, limiting both generalization and efficiency. We propose State of Thought (SoT), a new reasoning paradigm that enables endogenous reasoning in LLMs, with the model's internal reasoning state governing how reasoning unfolds. Concretely, SoT extracts a compact dynamics-geometric state from the model's internal information transfer and uses a 582-parameter controller on frozen backbones to selectively activate historical reasoning support useful under the current reasoning state, framing reasoning as a state-conditioned process over evidence rather than an externally prescribed token chain. Across quantitative (1.34x), general (1.62x), symbolic-and-code (1.76x), and long-context (2.51x) reasoning on 3 LLMs and 16 datasets, SoT consistently improves mean-baseline accuracy while reducing generated tokens by 62.6% and end-to-end latency by 44.6%. Across 2 VLM scales and 3 reasoning tasks, it improves mean accuracy by 3.8 points over reasoning baselines, with 74.9% fewer completion tokens and 73.5% lower latency than search-based methods. Under constrained access, SoT retains 38.2%/36.5% mean accuracy gains in training-free/embedding-only settings, while trajectory-only judging reaches 84.1% agreement across 3 API models. Together, endogenous state-driven reasoning provides a generalizable and efficient alternative.
Sources
- Program Synthesis with Large Language Models
- Qwen2.5-VL Technical Report
- Probing the Trajectories of Reasoning Traces in Large Language Models
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Training Verifiers to Solve Math Word Problems
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
- Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
- The Llama 3 Herd of Models
- Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
- Early Stopping for Large Reasoning Models via Confidence Dynamics
- Mixtral of Experts
- Think in Sentences: Explicit Sentence Boundaries Enhance Language Model's Capabilities
- Self-Refine: Iterative Refinement with Self-Feedback
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering