Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference
cs.LG, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: Accepted by Transactions on Machine Learning Research (TMLR), 2026
Journal ref: Transactions on Machine Learning Research, 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Greedy decoding from large language models is commonly treated as deterministic.
Terminology
Abstract
Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100% of prompts diverge; a single token flip often cascades into trajectory-level divergence. We develop an empirical error-propagation analysis and find that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. The analysis makes five testable predictions about intervention outcomes, including that applying more FP32 compute (broader scope) makes agreement worse. The experiments match all five predictions. The best-performing low-overhead intervention we evaluate, selective FP32 LM head recomputation, triggered only when the margin falls below a threshold, delivers +22-36 pp exact agreement on A10G (+12-21 pp on L4 and A100) at less than 4% latency overhead in low-batch (batch size <=4) single-stream inference. We map the applicability boundary across six models and four batch sizes, and hypothesise that training-time precision stability is a determining factor. The method is a partial mitigation rather than a universal determinism guarantee: its benefit vanishes when body-originated error dominates, including at batch size >=8 and under end-to-end FP8 in our tests.
Sources
- Non-Determinism of "Deterministic" LLM Settings
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism
- Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models
- LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
- A Study of BFLOAT16 for Deep Learning Training
- Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
- OLMoE: Open Mixture-of-Experts Language Models
- Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
- Defeating the Training-Inference Mismatch via FP16
- Qwen2.5 Technical Report
- Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements
- TinyLlama: An Open-Source Small Language Model
- Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks