Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs

arXiv:2509.13813 · cs.CL, cs.LG · Submitted 2025-09-17 · Read on arXiv

cs.CL, cs.LG

Submitted: 2025-09-17

Updated: 2026-09-22

Comments: 24 pages, 8 figures. Camera-ready version, published in Transactions on Machine Learning Research (2026). OpenReview: https://openreview.net/forum?id=5UVv7gkgUD

Journal ref: Transactions on Machine Learning Research (2026)

Code: https://github.com/zlin7/UQ-NLG

License: http://creativecommons.org/licenses/by/4.0/

The gist: Large language models are known to hallucinate, generating linguistically plausible but incorrect answers to questions.

Terminology

Abstract

Large language models are known to hallucinate, generating linguistically plausible but incorrect answers to questions. Uncertainty quantification has been proposed as a strategy to detect such behaviour, but existing methods lack a unified framework to assess reliability at both the prompt and answer level. We introduce a geometric framework which quantifies language model uncertainty at both levels by explicitly modelling a prompt-conditioned semantic distribution in answer embedding space. Our approach is black-box and sampling-based; we generate multiple answers per prompt, and use archetypal analysis to estimate a geometric support for the answer distribution. At the prompt level, we approximate the distribution entropy to quantify uncertainty; for each individual answer, we then use notions of atypicality to assess its reliability relative to the batch. We employ our framework to not only detect hallucinations but correct them, by selecting the batch example deemed most reliable. Experiments show that our framework performs comparably to or better than prior methods on short form question-answering datasets, and achieves superior results on medical datasets where hallucinations carry particularly critical risks. Beyond pure performance, we suggest the theoretical grounding of our work provides support for semantic distributions as useful objects of study for language model uncertainty.

Sources

Related papers