The Score Granularity Gap in Black-Box LLM Classification: A Comparative Study of Confidence Constructions
cs.CL
Submitted: 2026-06-20
Updated: 2026-08-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) are increasingly deployed as black-box classifiers in pipelines that automate confident decisions and route uncertain ones to human review.
Terminology
Abstract
Large language models (LLMs) are increasingly deployed as black-box classifiers in pipelines that automate confident decisions and route uncertain ones to human review. Such selective prediction needs a confidence score that an operator can threshold at a chosen risk level. Prior work asks whether LLM confidence is well calibrated or well ranked; we ask a complementary, deployment-oriented question that has been largely overlooked: at what resolution can the score be thresholded? We call the answer the score granularity gap. Through a controlled comparison of seven ways to build a confidence score, from a single verbalized number, to token probabilities, to querying the model many times and combining the answers, across 25 model-dataset pairs (9 LLMs, 3 benchmarks), we find that single-shot verbalized confidence, once correctly converted to a class probability, ranks cases surprisingly well, yet takes only a handful of distinct values. It therefore offers an operator only a few coarse thresholds, no matter how well it ranks. We show which constructions widen this gap, at what inference cost, and with what effect on ranking, notably that multi-query aggregation helps weak models but can degrade already-strong ones. We translate these trade-offs into concrete deployment guidance.
Sources
- Conformal Risk Control
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Selective Classification for Deep Neural Networks
- On Calibration of Modern Neural Networks
- PubMedQA: A Dataset for Biomedical Research Question Answering
- Language Models (Mostly) Know What They Know
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Teaching Models to Express Their Uncertainty in Words
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- Large Language Model Confidence Estimation via Black-Box Access
- Calibration in Deep Learning: A Survey of the State-of-the-Art
- Multi-Perspective Consistency Enhances Confidence Estimation in Large Language Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models
- Trust in One Round: Confidence Estimation for Large Language Models via Structural Signals
- Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering