Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation
cs.CL, cs.AI
Submitted: 2024-12-10
Updated: 2026-09-16
Comments: 8 pages, 5 figures. Accepted at IJCNN 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
- LM vs LM: Detecting Factual Errors via Cross Examination
- A Survey on In-context Learning
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
- Enhancing Confidence Expression in Large Language Models Through Learning from Past Experience
- Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
- Mistral 7B
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Language Models (Mostly) Know What They Know
- Scaling Laws for Neural Language Models
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Teaching Models to Express Their Uncertainty in Words
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
- Uncertainty Estimation in Autoregressive Structured Prediction
- Correcting Length Bias in Neural Machine Translation
- Are NLP Models really able to Solve Simple Math Word Problems?
- The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering