LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
cs.CL, cs.AI, cs.LG
Submitted: 2025-06-22
Updated: 2026-09-12
Comments: Codebase: https://github.com/yangalan123/LLMBranchingFactor. V4: TMLR 2026 Accepted. See https://openreview.net/forum?id=KotVuXj6CL¬eId=lGqX4g8e5V for what's changed
Code: https://github.com/yangalan123/LLMBranchingFactor
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite their impressive capabilities, aligned large language models (LLMs) often generate outputs that lack diversity.
Terminology
Abstract
Despite their impressive capabilities, aligned large language models (LLMs) often generate outputs that lack diversity. What drives this consistency in the generation? We investigate this phenomenon through the lens of probability concentration in the model's output distribution. To quantify it, we use the Branching Factor (BF)--the exponentiated length-averaged entropy of the output distribution, interpreted as the effective number of plausible next steps during generation. Our empirical analysis reveals two key findings: (1) BF often decreases as generation progresses, suggesting that LLMs become more predictable; a controlled intervention indicates that this decline is largely a task-independent property of autoregressive self-conditioning, distinct from alignment, and can be locally reversed by unexpected context. (2) Alignment tuning sharpens the output distribution from the outset, reducing BF by a factor of 2--5 overall and up to an order of magnitude (e.g., from 12 to 1.2) at early positions. This reduction helps explain why aligned models are less sensitive to decoding strategies. It also has implications for complex reasoning: aligned Chain-of-Thought (CoT) models (e.g., DeepSeek-distilled models) generate longer reasoning chains that reach later, more deterministic (lower-BF) stages, yielding more stable outputs. We hypothesize that alignment does not fundamentally change model behavior, but instead steers the model toward stylistic tokens (e.g., ``Sure'') that unlock low-entropy trajectories already present in the base model. Nudging experiments support this view: prompting base models with such tokens similarly reduces BF. Together, our findings establish BF as a diagnostic for understanding and controlling LLM outputs, clarifying how alignment reduces variability, CoT stabilizes generation, and base models can be steered away from diversity.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Reasoning with Exploration: An Entropy Perspective
- Modifying Large Language Model Post-Training for Diverse Creative Writing
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Calibration of Pre-trained Transformers
- The Llama 3 Herd of Models
- Nudging: Inference-time Alignment of LLMs via Guided Decoding
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Benchmarking Linguistic Diversity of Large Language Models
- Creative Preference Optimization
- Language Models (Mostly) Know What They Know
- GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets
- Diverse Preference Optimization
- Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
- Revisiting the Uniform Information Density Hypothesis
- Slaves to the Law of Large Numbers: An Asymptotic Equipartition Property for Perplexity in Generative Language Models
- 2 OLMo 2 Furious
- Qwen3 Technical Report
- The Effect of Sampling Temperature on Problem Solving in Large Language Models
- A Statistical Case Against Empirical Human-AI Alignment
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering