Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation
cs.CL
Submitted: 2026-01-22
Updated: 2026-09-10
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The challenge of uncertainty quantification of large language models in medicine
- Uncertainty Quantification for Clinical Outcome Predictions with (Large) Language Models
- Rainproof: An Umbrella To Shield Text Generators From Out-Of-Distribution Data
- Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Teaching Models to Express Their Uncertainty in Words
- Language Models (Mostly) Know What They Know
- Uncertainty Estimation of Large Language Models in Medical Question Answering
- Enhancing Healthcare LLM Trust with Atypical Presentations Recalibration
- The Llama 3 Herd of Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering