Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA
cs.CL, cs.LG
Submitted: 2026-01-30
Updated: 2026-09-18
Comments: Accepted to the Third Workshop on Uncertainty-Aware NLP at EMNLP 2026
Code: https://github.com/muelphil/llm-uncertainty-bench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Linguistic Calibration of Long-Form Generations
- From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
- LLMs Will Always Hallucinate, and We Need to Live With This
- A Diachronic Perspective on User Trust in AI under Uncertainty
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
- Uncertainty Toolbox: an Open-Source Library for Assessing, Visualizing, and Improving Uncertainty Quantification
- Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Unsupervised Quality Estimation for Neural Machine Translation
- Evaluating language models as risk scores
- Survey of Hallucination in Natural Language Generation
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Language Models (Mostly) Know What They Know
- Enhancing Confidence Expression in Large Language Models Through Learning from Past Experience
- A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
- Large Language Models Must Be Taught to Know What They Don't Know
- What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
- Measuring Massive Multitask Language Understanding
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Do LLMs estimate uncertainty well in instruction-following?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering