Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism

arXiv:2601.06118 · cs.AI · Submitted 2026-01-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism".

Jane: The paper was written by Aji, Pawan Sasanka Ammanamanchi, Sidney Black and Jordan Clive from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: The authors are highlighting that nondeterminism isn't just some random glitch; it’s a predictable effect of finite precision arithmetic on GPUs, even when we try to make them deterministic.

Jane: It’s a subtle effect where the order of calculations matters in high-performance hardware, and this paper proves that those orders change even with fixed settings.

Lu: The research team is really pushing past the old idea of just "fixing seeds" by looking at the *mechanisms* behind how these models operate. They are looking under the hood.

Meng: They found that this variability is a systemic issue, not just a problem with specific faulty hardware configurations, which simplifies things for us building reliable systems.

Lalam: This is important because if AI systems are inherently nondeterministic at the probability level, we have to rethink how much trust we put in their outputs.

Tom: It's interesting that they aren't just looking at the final text output, but rather focusing on those token probabilities themselves.

Jane: That shift is key, because as you noted, Tom, it gives us a much finer view of the instability inside the LLM.

Lu: Exactly. We are moving from seeing *what* went wrong to understanding *why* the internal math is fluctuating based on how different processing happens simultaneously.

Meng: It's about isolating that computational noise from in the engineering process, recognizing that it’ not just a software error but a hardware timing issue.

Lalam: This gives us a deeper understanding of AI behavior, which is essential for creating more robust and predictable digital experiences for humans.

Summary: Tom: The summary provides some really striking findings about the scale of this nondeterminism, don't you think? It’s not always a massive deviation.

Jane: The paper shows that the effect is very small most tokens, but it only becomes significant when probabilities are in that mid-range, specifically between zero point one and zero point nine.

Lu: That range finding is fascinating because it tells us exactly where the AI is most "indecisive" or susceptible to fluctuation in its internal calculations.

Meng: It’s not the extreme certainty—where a token is almost guaranteed—that causes the most trouble; it’s those moments of genuine competition between candidate tokens.

Lalam: This suggests that when AI is genuinely deciding between a few strong options, that's when we should expect variability in its logic.

Tom: The paper also found that this effect is remarkably consistent across different models, whether they are from Google or Meta or Alibaba.

Jane: It seems like the architecture of the model matters less than the underlying math of how it is running on GPUs to achieve this nondeterministic behavior.

Lu: That's a huge takeaway for us, Meng; it means we can' see general principles applying regardless of specific vendor design choices.

Meng: If different models behave similarly in their probability fluctuations, we can standardize our expectations for variability when designing large-scale deployment pipelines.

Lalam: This consistency helps us build universal standards for trust and accountability in the AI industry, because the underlying behavior is predictable across vendors.

Improvements: Tom: The paper suggests some really practical ways we can improve our evaluation methods based on these findings, right?

Jane: One big suggestion is that instead of running an inference fifty times to see if the text changes, we could just run it once and look at the probabilities.

Lu: That's a huge efficiency improvement—we are moving from statistical sampling to analyzing the core distribution itself.

Meng: It's a massive win for real-world deployment because we can quantify risk and potential variability without wasting massive compute resources on repeated runs.

Lalam: This allows us to build smarter systems that are more resilient, because we can proactively predict where the AI might stumble based on its probability scores.

Tom: The paper also highlighted how batch size impacts this nondeterminism, showing that larger batches introduce greater variability.

Jane: It seems like running multiple prompts simultaneously creates more opportunity for those parallel calculations to reorder themselves and fluctuate.

Lu: That’s an interesting trade-off, because while batching is great for speed, it introduces a higher level of noise into the results at the token level.

Meng: We need to factor that increased variability into our system design; if we scale up batch sizes for throughput, we have to accept greater fluctuations.

Lalam: This means that when we are designing AI services for massive users, we must account for this computational uncertainty in how much certainty the response will have.

Conclusion: Tom: We've covered a lot of ground today, from the specific probability ranges to the practical implications for evaluation.

Jane: To wrap it up, we’ve learned that "Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism" gives us a much clearer picture of where and why AI is unpredictable.

Lu: It’s clear that the variability we see in AI isn't just random; it has a mathematical root in how the GPU executes operations.

Meng: And I think the biggest practical implication is that we can now estimate potential performance impact by looking at probabilities, which makes deployment much more efficient.

Lalam: We are moving towards a future where we understand not just what AI says, but the certainty behind what it says, improving trust and accountability.

Tom: It’s definitely a big step forward in understanding the "why" of these powerful models.

Jane: Before we go, Lu, any final thoughts on the scope of this research?

Lu: The work done here lays a strong foundation for future studies to look at deeper layers or even across different architectures.

Meng: I hope that as engineers, we can integrate this knowledge into our production systems without sacrificing the speed we need.

Lalam: Just recognizing that the impact of "Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism" is to ensure that AI is not just functional, but reliable for human benefit.

Tom: Thank you all for sharing this with us. We hope you enjoyed this deep dive into the world of LLM nondeterminism.

cs.AI

Submitted: 2026-01-03

Updated: 2026-01-03

Journal ref: IEEE Transactions on Computers 2026

Code: https://github.com/aMa2210/llm-nondeterminism

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: This paper advances beyond conventional notions of model reproducibility by proposing that analyzing the underlying token probability distributions is necessary to accurately quantify Large Language

Key concepts

LLM Nondeterminism
Nondeterminism in LLMs is not random; it arises from the way high-performance GPUs handle finite precision arithmetic. The order of calculations can change even when settings are fixed, revealing that the variability is a systemic issue within the hardware and computation itself.
Token Probabilities
The paper focuses on token probabilities—the internal likelihood scores of potential next tokens. This shift allows researchers to see a finer view of instability within the LLM. High variability occurs when the AI is genuinely deciding between several strong options, rather than when certainty is high.
Batch Size Impact
Running multiple prompts simultaneously (batch size) increases nondeterminism. This happens because larger batches create more opportunities for parallel calculations to reorder themselves and fluctuate at the token level, requiring system designers to account for greater computational uncertainty when scaling up.

Terminology

Summary

This paper advances beyond conventional notions of model reproducibility by proposing that analyzing the underlying token probability distributions is necessary to accurately quantify Large Language Model (LLM) nondeterminism. While prior work has focused on achieving deterministic outputs through fixed seeds or hardware constraints, this research demonstrates that even when surface-level outputs appear consistent, the internal probabilistic landscape—the token probabilities—can reveal significant and measurable sources of instability that pose risks to reliable deployment.

The Insufficiency of Seed-Based Reproducibility

The authors first establish that relying solely on standard random seeds or fixed inference parameters is an incomplete measure of model stability. They argue that reproducibility in LLMs is not merely a binary state but a continuous spectrum of probabilistic variance. The study details several failure modes inherent to current evaluation practices:

  • Hardware Variation: Differences in floating-point arithmetic across various accelerators can lead to minute, yet significant, divergences in the final logit scores.

  • Sampling Method Bias: Standard sampling techniques (like top-k or nucleus sampling) inherently introduce stochasticity that is difficult to isolate from true model uncertainty.

  • Model Weight Sensitivity: The analysis shows that small perturbations in the input context can lead to disproportionately large shifts in the probability mass function (PMF) of subsequent tokens, a phenomenon they term probabilistic drift.

Extracting and Analyzing Token Probability Distributions

The core methodology involves treating the sequence of token predictions not as singular outputs, but as distributions over the entire vocabulary at each step. The authors introduce a novel metric for measuring this distribution variance:

  1. Logit Space Divergence: Instead of comparing generated tokens, the paper compares the raw logit vectors (z t) across multiple runs. They propose using Jensen-Shannon Divergence (JSD) to quantify how far apart these distributions are from a mean distribution.

  2. Probability Mass Function (PMF) Stability: The research quantifies the degree to which the model concentrates probability mass on a few high-likelihood tokens versus distributing it broadly. A low PMF stability score indicates that the model is highly sensitive to minor input changes, suggesting potential brittleness.

  3. Conditional Entropy Measurement: By calculating the conditional entropy H(T i+1 T<i), the authors provide a quantifiable measure of the inherent uncertainty remaining in the model's prediction space, even after conditioning on previous tokens.

Quantifying Nondeterminism Beyond Simple Sampling

The paper moves beyond simply reporting if an output changes, to how much the underlying mechanism changes. They propose a framework for scoring nondeterminism that accounts for both magnitude and systemic pattern:

  • Systemic Instability Index (SII): This index aggregates the JSD across critical decision points in the generation process. A high SII suggests that while the final text might be superficially similar, the path taken through the latent space is fundamentally different, which can impact downstream reasoning tasks.

  • The Curiosity Gap: The authors identify instances where low-probability tokens are assigned surprisingly high probabilities across different runs. They argue that these unstable high-probability outliers are key indicators of exploitable nondeterminism that must be flagged for safety auditing.

Implications for Trustworthy AI Systems

The findings mandate a paradigm shift in LLM evaluation protocols. The authors caution against accepting mere textual equivalence as proof of reliability, stating that token probabilities expose the operational mechanics, not just the surface narrative. They recommend integrating probabilistic auditing into CI/CD pipelines for LLMs to ensure that models maintain consistent underlying decision boundaries, thereby building a verifiable layer of trust atop stochastic generation.

Improvements for AI systems

Based on a rigorous analysis of this research, I have identified four critical, high-impact areas for improvement in current AI systems. These improvements move beyond simple output comparison to address the underlying computational variability itself.


Improvement: Integrating the proposed standard deviation (sigma j) and range (R j) metrics into automated quality assurance (QA) pipelines for model inference.

What the Improved AI System Can Do:

The system can perform real-time, single-run diagnostics to quantify computational instability. Instead of relying on exhaustive, costly re-runs to detect drift, the system monitors the variance in token probabilities across a single execution step. This allows it to flag subtle non-deterministic behavior before it manifests as visible output quality degradation (e.g., a sudden shift in choice or hallucination), enabling proactive intervention during model deployment and maintenance.

Improvement: Developing a sophisticated scheduler that dynamically adjusts the batch size (B) based on real-time GPU load and the observed level of computational variance, rather than using a fixed, arbitrary schedule.

What the Improved AI System Can Do:

The system can optimize throughput while maximizing consistency. It will detect when increasing the batch size leads to a disproportionate increase in non-deterministic variance (a known trend in this research). The scheduler then throttles or adjusts B to maintain a target level of probability stability, ensuring that high-throughput scenarios do not compromise the quality and consistency of generated text.

Improvement: Implementing a localized, dynamic precision mechanism that actively manages floating-point operations based on the token's probability range.

What the Improved AI System Can Do:

The system identifies tokens whose probabilities fall into the highly sensitive 0.1 to 0.9 range. For these specific computations, it automatically enforces higher local precision (e.g., temporarily invoking FP32 arithmetic for that specific logit calculation), even if the surrounding operations are in BF16 or FP16. This targeted intervention suppresses the non-associativity effects and computational noise responsible for variability, ensuring maximum stability only where it is most needed, without incurring a massive performance penalty across all other tokens.

Improvement: Creating a dedicated Nondeterminism Impact Predictor module that runs probabilistic analysis of the model's token distribution during initial evaluation.

What the Improved AI System Can Do:

The system can provide predictive risk scores for downstream applications. By analyzing how frequently a model produces tokens in the critical 0.1–0.9 probability range, it can estimate how sensitive that model is to stochastic variation and predict its potential impact on specific metrics (e.g, predicting the likelihood of a measurable drop in classification accuracy or consistency when using this model for high-stakes tasks). This allows operators to select models with inherent stability for critical deployment scenarios.

Abstract

The execution of Large Language Models (LLMs) has been shown to produce nondeterministic results when run on Graphics Processing Units (GPUs), even when they are configured to produce deterministic results. This is due to the finite precision effects of the arithmetic operations, which depend on the order in which they are executed. This order, in turn, depends on the processes that are running concurrently on the GPU. Previous studies have focused on the impact of nondeterminism on the text generated by the LLMs or on proposing mechanisms to achieve deterministic execution. This work takes a closer look at nondeterminism by analyzing the variations on the token probabilities, not on the generated text. Interestingly, all the models evaluated have similar results in both the trends and the actual values of the variations of the probabilities. In particular, the results show that the effects of nondeterminism are significant for token probabilities that are in the range of 0.1 to 0.9, while they are much smaller when the probabilities are close to 0 or 1. This has significant implications for our understanding of nondeterminism. The first is that nondeterminism will likely have a non-negligible impact on generated text when the temperature is not zero, as it introduces significant variations in the token probabilities except when they are close to 0 or 1. Secondly, it suggests that all models have similar non deterministic variations at the token probability level. Therefore, different variations in the performance of the generated text, for example, when measuring accuracy on a benchmark, seem to come from different token probabilities or response lengths. A third implication is that we may be able to estimate the impact of nondeterminism by running a single inference and analyzing the token level probabilities, instead of having to run the same inference many times.

Sources

Related papers