Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning
cs.AI, cs.CL, cs.IT, cs.LG, math.IT
Submitted: 2026-09-08
Updated: 2026-10-05
Comments: 16 pages, 4 figures, 3 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chain-of-thought reasoning provides a structured computation between a model's input and final answer.
Terminology
Abstract
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves over the reasoning process but do not reveal which competing hypotheses account for that uncertainty. We introduce answer-distribution trajectories, a stochastic-dynamics-inspired representation that tracks the model's full predictive distribution over answers as reasoning unfolds. As a strictly finer representation than endpoint and entropy summaries, answer-distribution trajectories enable us to characterize a trace through a dynamical reasoning profile spanning exploration, revision, motion, and commitment, and to distinguish different dynamical mechanisms of reasoning success and failure. Across sixteen open-weight language models and four reasoning benchmarks, we show that traces with the same endpoint and similar entropy profiles can exhibit substantially different reasoning dynamics. We further find substantial variation in these dynamics both within and across models and tasks, with different objectives favoring different dynamical profiles. Additionally, we show that training and inference choices systematically reshape these profiles. Our results suggest that answer-distribution trajectories provide a rich framework for analysing and evaluating the dynamics of LLM reasoning.
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models
- The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
- INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Early Stopping for Large Reasoning Models via Confidence Dynamics
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Measuring Faithfulness in Chain-of-Thought Reasoning
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- OpenAI o1 System Card
- Are NLP Models really able to Solve Simple Math Word Problems?
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- Qwen2.5 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection