Temperature Fragility and the Conditional Benefits of Truncation Sampling
cs.CL, cs.LG
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 20 pages, 8 figures, 16 tables. Code and data: https://github.com/larosafrancesco289/decoding-robustness
Code: https://github.com/larosafrancesco289/decoding-robustness
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models generate text by sampling each token from a predicted distribution, and a temperature parameter sets how far the draw strays from the most probable tokens.
Terminology
Abstract
Large language models generate text by sampling each token from a predicted distribution, and a temperature parameter sets how far the draw strays from the most probable tokens. Truncation samplers such as top-p and min-p discard the least probable tokens before the draw, so that sampling at high temperature stays coherent. Their reported accuracy gains come from temperatures of 1.5 to 3, while the defaults of deployed systems cluster between 0.6 and 1.0. Whether they change accuracy at those defaults, and for which models, has not been measured. We test thirteen open-weight models on GSM8K and MMLU-Pro at temperatures 0.7, 1.0, and 1.3 in one controlled pipeline, ten of them under eight decoding configurations. Six of the thirteen models lose 17 to 38 accuracy points on MMLU-Pro between 0.7 and 1.3, and the other seven lose at most 10. The lost accuracy comes from generations that run to the token limit or never state an answer. These results suggest that truncation samplers improve accuracy primarily when higher temperatures substantially degrade model performance. Where accuracy remains stable across temperatures, none of the tested truncation samplers improves on plain temperature sampling.
Sources
- From Confidence to Collapse in LLM Factual Robustness
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Reliability Under Randomness: An Empirical Analysis of Sparse and Dense Language Models Across Decoding Temperatures
- A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
- Training Verifiers to Solve Math Word Problems
- Mistral 7B
- What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend
- The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models
- Top-$n\sigma$: Not All Logits Are You Need
- Hermes 3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering