Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

arXiv:2609.25624 · cs.AR, cs.LG · Submitted 2026-09-22 · Read on arXiv

cs.AR, cs.LG

Submitted: 2026-09-22

Updated: 2026-09-22

Comments: 14 pages, 5 figures

Code: https://github.com/lpc97667/rf

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers