Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
cs.CL, cs.AI, cs.LG
Submitted: 2026-05-27
Updated: 2026-08-26
Code: https://github.com/cvs-health/uqlm
Terminology
Sources
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
- MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation
- Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness
- Selectively Answering Ambiguous Questions
- RED-CT: A Systems Design Methodology for Using LLM-labeled Data to Train and Deploy Edge Classifiers for Computational Social Science
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- Internal Representations as Indicators of Hallucinations in Agent Tool Selection
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
- Language Models (Mostly) Know What They Know
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
- Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
- Uncertainty Estimation in Autoregressive Structured Prediction
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
- Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic Space
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering