Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism
summary
The gist
This paper advances beyond conventional notions of model reproducibility by proposing that analyzing the underlying token probability distributions is necessary to accurately quantify Large Language
In short
The episode discusses a paper demonstrating that Large Language Model (LLM) nondeterminism stems from finite precision arithmetic on GPUs, not random glitches. This variability is systemic and affects token probabilities most when they are between 0.1 and 0.9. A key practical takeaway is using probability analysis instead of repeated inference runs to evaluate risk efficiently.
Key concepts
- LLM Nondeterminism
- Nondeterminism in LLMs is not random; it arises from the way high-performance GPUs handle finite precision arithmetic. The order of calculations can change even when settings are fixed, revealing that the variability is a systemic issue within the hardware and computation itself.
- Token Probabilities
- The paper focuses on token probabilities—the internal likelihood scores of potential next tokens. This shift allows researchers to see a finer view of instability within the LLM. High variability occurs when the AI is genuinely deciding between several strong options, rather than when certainty is high.
- Batch Size Impact
- Running multiple prompts simultaneously (batch size) increases nondeterminism. This happens because larger batches create more opportunities for parallel calculations to reorder themselves and fluctuate at the token level, requiring system designers to account for greater computational uncertainty when scaling up.
Terminology used across episodes
This episode discusses
- Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism · Paper Radio
- Non-Determinism of "Deterministic" LLM Settings
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
- Language Models are Few-Shot Learners
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
- Measuring Massive Multitask Language Understanding
- The Curious Case of Neural Text Degeneration
- Generative Artificial Intelligence Reproducibility and Consensus
- Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
- FP8 Formats for Deep Learning
- GPT-4 Technical Report
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference
The paper
Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism · Read on arXiv
The execution of Large Language Models (LLMs) has been shown to produce nondeterministic results when run on Graphics Processing Units (GPUs), even when they are configured to produce deterministic results. This is due to the finite precision effects of the arithmetic operations, which depend on the order in which they are executed. This order, in turn, depends on the processes that are running concurrently on the GPU. Previous studies have focused on the impact of nondeterminism on the text generated by the LLMs or on proposing mechanisms to achieve deterministic execution. This work takes a closer look at nondeterminism by analyzing the variations on the token probabilities, not on the generated text. Interestingly, all the models evaluated have similar results in both the trends and the actual values of the variations of the probabilities. In particular, the results show that the effects of nondeterminism are significant for token probabilities that are in the range of 0.1 to 0.9, while they are much smaller when the probabilities are close to 0 or 1. This has significant implications for our understanding of nondeterminism. The first is that nondeterminism will likely have a non-negligible impact on generated text when the temperature is not zero, as it introduces significant variations in the token probabilities except when they are close to 0 or 1. Secondly, it suggests that all models have similar non deterministic variations at the token probability level. Therefore, different variations in the performance of the generated text, for example, when measuring accuracy on a benchmark, seem to come from different token probabilities or response lengths. A third implication is that we may be able to estimate the impact of nondeterminism by running a single inference and analyzing the token level probabilities, instead of having to run the same inference many times.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism".
Jane: The paper was written by Aji, Pawan Sasanka Ammanamanchi, Sidney Black and Jordan Clive from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: The authors are highlighting that nondeterminism isn't just some random glitch; it’s a predictable effect of finite precision arithmetic on GPUs, even when we try to make them deterministic.
Jane: It’s a subtle effect where the order of calculations matters in high-performance hardware, and this paper proves that those orders change even with fixed settings.
Lu: The research team is really pushing past the old idea of just "fixing seeds" by looking at the *mechanisms* behind how these models operate. They are looking under the hood.
Meng: They found that this variability is a systemic issue, not just a problem with specific faulty hardware configurations, which simplifies things for us building reliable systems.
Lalam: This is important because if AI systems are inherently nondeterministic at the probability level, we have to rethink how much trust we put in their outputs.
Tom: It's interesting that they aren't just looking at the final text output, but rather focusing on those token probabilities themselves.
Jane: That shift is key, because as you noted, Tom, it gives us a much finer view of the instability inside the LLM.
Lu: Exactly. We are moving from seeing *what* went wrong to understanding *why* the internal math is fluctuating based on how different processing happens simultaneously.
Meng: It's about isolating that computational noise from in the engineering process, recognizing that it’ not just a software error but a hardware timing issue.
Lalam: This gives us a deeper understanding of AI behavior, which is essential for creating more robust and predictable digital experiences for humans.
Summary: Tom: The summary provides some really striking findings about the scale of this nondeterminism, don't you think? It’s not always a massive deviation.
Jane: The paper shows that the effect is very small most tokens, but it only becomes significant when probabilities are in that mid-range, specifically between zero point one and zero point nine.
Lu: That range finding is fascinating because it tells us exactly where the AI is most "indecisive" or susceptible to fluctuation in its internal calculations.
Meng: It’s not the extreme certainty—where a token is almost guaranteed—that causes the most trouble; it’s those moments of genuine competition between candidate tokens.
Lalam: This suggests that when AI is genuinely deciding between a few strong options, that's when we should expect variability in its logic.
Tom: The paper also found that this effect is remarkably consistent across different models, whether they are from Google or Meta or Alibaba.
Jane: It seems like the architecture of the model matters less than the underlying math of how it is running on GPUs to achieve this nondeterministic behavior.
Lu: That's a huge takeaway for us, Meng; it means we can' see general principles applying regardless of specific vendor design choices.
Meng: If different models behave similarly in their probability fluctuations, we can standardize our expectations for variability when designing large-scale deployment pipelines.
Lalam: This consistency helps us build universal standards for trust and accountability in the AI industry, because the underlying behavior is predictable across vendors.
Improvements: Tom: The paper suggests some really practical ways we can improve our evaluation methods based on these findings, right?
Jane: One big suggestion is that instead of running an inference fifty times to see if the text changes, we could just run it once and look at the probabilities.
Lu: That's a huge efficiency improvement—we are moving from statistical sampling to analyzing the core distribution itself.
Meng: It's a massive win for real-world deployment because we can quantify risk and potential variability without wasting massive compute resources on repeated runs.
Lalam: This allows us to build smarter systems that are more resilient, because we can proactively predict where the AI might stumble based on its probability scores.
Tom: The paper also highlighted how batch size impacts this nondeterminism, showing that larger batches introduce greater variability.
Jane: It seems like running multiple prompts simultaneously creates more opportunity for those parallel calculations to reorder themselves and fluctuate.
Lu: That’s an interesting trade-off, because while batching is great for speed, it introduces a higher level of noise into the results at the token level.
Meng: We need to factor that increased variability into our system design; if we scale up batch sizes for throughput, we have to accept greater fluctuations.
Lalam: This means that when we are designing AI services for massive users, we must account for this computational uncertainty in how much certainty the response will have.
Conclusion: Tom: We've covered a lot of ground today, from the specific probability ranges to the practical implications for evaluation.
Jane: To wrap it up, we’ve learned that "Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism" gives us a much clearer picture of where and why AI is unpredictable.
Lu: It’s clear that the variability we see in AI isn't just random; it has a mathematical root in how the GPU executes operations.
Meng: And I think the biggest practical implication is that we can now estimate potential performance impact by looking at probabilities, which makes deployment much more efficient.
Lalam: We are moving towards a future where we understand not just what AI says, but the certainty behind what it says, improving trust and accountability.
Tom: It’s definitely a big step forward in understanding the "why" of these powerful models.
Jane: Before we go, Lu, any final thoughts on the scope of this research?
Lu: The work done here lays a strong foundation for future studies to look at deeper layers or even across different architectures.
Meng: I hope that as engineers, we can integrate this knowledge into our production systems without sacrificing the speed we need.
Lalam: Just recognizing that the impact of "Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism" is to ensure that AI is not just functional, but reliable for human benefit.
Tom: Thank you all for sharing this with us. We hope you enjoyed this deep dive into the world of LLM nondeterminism.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language