Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA

arXiv:2602.00279 · cs.CL, cs.LG · Submitted 2026-01-30 · Read on arXiv

cs.CL, cs.LG

Submitted: 2026-01-30

Updated: 2026-09-18

Comments: Accepted to the Third Workshop on Uncertainty-Aware NLP at EMNLP 2026

Code: https://github.com/muelphil/llm-uncertainty-bench

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers