Semantic Self-Distillation for Language Model Uncertainty
cs.CL, cs.LG
Submitted: 2026-02-04
Updated: 2026-09-22
Comments: Camera-ready version, published in Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026), PMLR 337:5427-5447
Journal ref: Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:5427-5447, 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models present challenges for principled uncertainty quantification, in part due to their complexity and the diversity of their outputs.
Terminology
Abstract
Large language models present challenges for principled uncertainty quantification, in part due to their complexity and the diversity of their outputs. Semantic dispersion, or the variance in the meaning of sampled answers, has been proposed as a useful proxy for model uncertainty, but the associated computational cost prohibits its use in latency-critical applications. We show that sampled semantic distributions can be distilled into lightweight student models which estimate a prompt-conditioned density before the language model generates an answer token. The student model predicts a semantic distribution over possible answers; the entropy of this distribution provides a prompt-level uncertainty signal, and the probability density allows answer-level reliability evaluation. Across experiments on TriviaQA and MMLU, we find our student models perform competitively relative to the teacher's sampled semantic dispersion on a hallucination prediction task, whilst offering additional uncertainty primitives for out-of-domain detection and multiple-choice answer selection. We term this technique Semantic Self-Distillation (SSD), which can serve as a general framework for distilling predictive uncertainty in complex output spaces beyond language.
Sources
- The Llama 3 Herd of Models
- Language Models (Mostly) Know What They Know
- Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
- Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
- Real-Time Detection of Hallucinated Entities in Long-Form Generation
- Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs
- Gemma 3 Technical Report
- EmbeddingGemma: Powerful and Lightweight Text Representations
- Qwen3 Technical Report
- TrajSurv: Learning Continuous Latent Trajectories from Electronic Health Records for Trustworthy Survival Prediction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering