SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation".
Tom: Span-Level Uncertainty Estimation (SLUE) formalizes a new task targeting semantically coherent text spans to provide interpretable and localizable uncertainty scores for Large Language Model generations.
Jane: First, who's behind it and why it matters.
Title and authors: Jane: The authors introduce SPANUQ, which they describe as a lightweight probe with about twenty-five million parameters. This means it's designed to be efficient enough to run during inference without needing multiple expensive passes over the LLM hidden states.
Meng: That efficiency is huge for us; running something this small instead of running five or ten full sampling runs just to get an idea of uncertainty is a massive time and computational saving.
Lu: They achieve this by distilling the knowledge from those multi-sample inferences into a single forward pass over the LLM hidden states, which they detail in the SPANUQ: Span-Level Uncertainty Quantification for Large Language Model Generation paper. This distillation process is what makes it so practical for real-world use.
Tom: So, they are using this probe to get span-level uncertainty directly from the frozen LLM hidden states without needing complex external systems or heavy sampling during the main generation phase. That’s a really smart architectural move, I think.
Lalam: And that direct access to hidden states is powerful because it lets us inject this uncertainty estimation process right into the generation flow, rather than tacking it on afterwards as a separate step.
Jane: They use a DETR-style span decoder for detection and then pair that with a Mixture of Beta model for estimating the uncertainty parameters, which allows them to capture the bimodal nature of uncertainty in their SpanUQ paper.
Meng: The combination of those components sounds complex, but if it’s all distilled into one forward pass, I can see how it could fit into our existing inference infrastructure without a massive overhaul.
Tom: It’s not just that it’s efficient; the paper shows they get top performance on span-level quality metrics, achieving an AUROC between zero point nine zero eight and zero point nine four four on five different LLM backbones tested in their work, which is quite high for this kind of task.
Lu: And what really stands out to me from the SPANUQ: Span-Level Uncertainty Quantification for Large Language Model Generation paper is the span-to-sequence decomposability they observe, where the learned importance-weighted span composition achieves a rho seq of zero point eight three nine, suggesting that span estimation actually covers sequence estimation as a special case.
Lalam: That decomposition finding is huge because it gives us a way to understand exactly how much each part of the AI output contributes to the final reliability score, which is incredibly useful for auditing and trust assessment.
The paper's summary: Tom: So, if we look at what they summarize in "SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation," they are essentially proposing a new formal task called SLUE—Span-Level Uncertainty Estimation—which focuses on those semantically coherent text spans.
Jane: They argue that this task is the natural unit for uncertainty because each span conveys exactly one assessable piece of information, making the resulting score directly interpretable and actionable in a way token scores simply aren't.
Meng: It’s about moving from measuring the whole response to measuring individual, meaningful segments of it, which addresses that localization problem we talked about earlier when sequence-level methods fail to pinpoint where an error lies.
Lu: Page one explains this contrast well; they show how token-level scores are noisy and hard to interpret, whereas span-level uncertainty provides a continuous score between zero and one for each segment, which is much more informative.
Tom: And the paper details their framework: it uses multi-layer fusion of hidden states, projects them through a token encoder to get a content pool Z, and then uses DETR queries to detect spans that attend over that sequence information.
Jane: The feature enrichment module is where they make it really clever; they use a differentiable soft boundary mask and mask-weighted attention pooling to create enriched queries that bridge the gap between where the span is located and what content it actually represents.
Lalam: That enrichment step, using the gated residual connection, seems key because it’s how they manage to connect the location of a detected span with its actual semantic content for better uncertainty estimation.
Tom: And finally, they have prediction heads that do boundary regression and validity classification alongside their Mixture of Beta uncertainty estimation model to capture that bimodal nature of the uncertainty distribution.
Meng: So, in short, they take expensive multi-sample inference knowledge and compress it into a single forward pass that simultaneously detects spans and assigns those continuous uncertainty scores. That's the big technical summary for SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation.
The paper's improvements: Jane: The paper outlines several improvements to existing methods, focusing on overcoming the limitations of prior granularity levels, which are what they call token-level and sequence-level scores.
Tom: They specifically improve things by making the uncertainty estimation interpretable; instead of just a single score for a whole text, they provide localizable scores that tell you exactly which part is unreliable.
Lu: The key improvement is the introduction of span-to-sequence decomposability, which shows that span-level estimation subsumes sequence-level estimation as a special case because the learned importance weights are quite high.
Meng: From an engineering standpoint, this means we can build systems that don't just flag "this whole answer is bad," but can actually flag the specific sentence or clause within it that is causing the issue, which dramatically improves debugging workflows.
Lalam: That ability to decompose uncertainty into granular spans allows us to create self-refining LLMs; if we know exactly where the weakness is, we can target our refinement efforts precisely there instead of retraining everything.
Tom: They also introduce Uncertainty-Conditioned Iterative Refinement, or UCIR, at inference time, which feeds the initial estimates back into a second pass where queries are refined using an MLP adapter.
Jane: This iterative refinement step is significant because it corrects systematic errors that the first pass might have made, and they achieve this with only a small inference overhead of less than fifteen percent when using alpha=zero point seven.
Lu: So, the framework moves beyond just detection; it’s an improvement because it includes a refinement loop that actively tries to improve the initial uncertainty estimates based on subsequent passes.
Conclusion: Tom: So, to wrap up our discussion on "SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation," the main implication is that we have a method that moves beyond vague confidence scores by providing highly granular, localizable uncertainty scores for LLM output.
Jane: It means we can finally start performing true trust assessment on generated text with high precision, allowing users to know exactly where to focus their fact-checking efforts instead of wading through the entire response blindly.
Meng: For practical application, this is huge because it leads directly to more robust deployment in high-stakes environments like legal or medical fields where a calibrated confidence score that distinguishes between correct and potentially dangerous claims is essential.
Lalam: I think the most significant cultural impact is enabling self-refining LLMs, where the model can improve its own outputs based on this uncertainty feedback, leading to more coherent and grounded generations over time.
Lu: And from a research view, showing that span-level estimation subsumes sequence-level estimation proves that this approach is fundamentally sound and provides a new way to think about how we structure knowledge representation in language generation.
Tom: It’s been really fascinating dissecting the mechanics of SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation. We've seen how this probe can distill complex inference into a single pass that delivers incredibly detailed insights into model reliability.
Jane: I agree, Tom; it gives us a concrete tool to evaluate and guide the next generation of AI models with much more precision than we had before.
Meng: It’s definitely a framework we should be looking at seriously for improving our quality assurance pipelines moving forward because it offers measurable risk scores tied directly to content structure.
Lalam: I'm genuinely excited about seeing how this feeds into a self-refining system; that level of internal improvement is what makes an AI truly useful in complex reasoning tasks.
Yimeng Zhang, Yingying Zhuang, Ziyi Wang, Yuxuan Lu, Pei Chen, Aman Gupta, Zhe Su, Ming Tan
Amazon
cs.CL
Submitted: 2026-07-07
Updated: 2026-09-29
Comments: Accepted by NeurIPS 2026. The project page is available at https://damon-demon.github.io/SpanUQ.html
Project page: https://damon-demon.github.io/SpanUQ.html
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Span-Level Uncertainty Estimation (SLUE) formalizes a new task targeting semantically coherent text spans to provide interpretable and localizable uncertainty scores for Large Language Model
Key concepts
- Span-Level Uncertainty Estimation (SLUE)
- This is a new task where the goal is to detect contiguous text segments (spans) that form a single unit of meaning and assign them a continuous uncertainty score. This method focuses on localizing errors semantically, unlike token-level scores or sequence-level scores.
- Span Definition
- A span is defined as a contiguous text segment that conveys one coherent unit of meaning. This definition ensures that each span carries exactly one assessable piece of information, allowing its uncertainty score to be directly interpreted and actionable for users.
- Mixture of Beta (MoB) Uncertainty Estimation Model
- This is the specific model used by SPANUQ's prediction heads to estimate uncertainty. It captures the 'bimodal nature' of uncertainty, meaning it can model situations where a span is either highly certain or highly uncertain, providing a richer statistical description than simple point estimates.
Terminology
Summary
Span-Level Uncertainty Estimation (SLUE) formalizes a new task targeting semantically coherent text spans to provide interpretable and localizable uncertainty scores for Large Language Model generations. This approach addresses the limitations of token-level scores lacking semantic coherence and sequence-level scores failing to localize errors, making it essential for trustworthy LLM deployment.
The gist
SPANUQ achieves the best span-level uncertainty quality (AUROC 0.908–0.944, MAE 0.110–0.129), outperforming the strongest probe baseline and all sampling-based methods while being 10–20× faster than those requiring multiple forward passes over LLM hidden states.
Problem Formulation and Motivation
The paper formalizes SLUE as a new task: given a single LLM forward pass, jointly detect spans and assign continuous uncertainty scores to each.
The motivation stems from the inadequacy of existing granularity levels: token-level methods are semantically incomplete,
while sequence-level methods produce a single score that cannot localize which part is unreliable.
Spans are defined as a contiguous text segment that conveys a single, coherent unit of meaning whose uncertainty can be independently assessed.
This makes span-level estimation the natural unit for uncertainty estimation: each carries exactly one assessable piece of information, making its uncertainty score directly interpretable and actionable.
SPANUQ Framework Architecture
SPANUQ is introduced as a lightweight probe (∼25M parameter) that distills the uncertainty knowledge from expensive multi-sample inference into a single forward pass over LLM hidden states.
The architecture consists of five components:
-
Multi-Layer Fusion: Fuses hidden states from selected layers (e.g., 20–26 for Qwen3-14B) via
mean fusion
to capture complementary signals across syntax and factual knowledge. -
Token Encoder: Projects the fused states to a lower dimension and contextualizes them through a 2-layer Transformer encoder, yielding the content pool Z.
-
DETR-Style Span Decoder: Uses
N=32 learnable span queries Q
that attend to the token sequence through a 3-layer Transformer decoder, enabling parallel detection of spans. -
Span Feature Enrichment: This module bridges the gap between location and content by using a
differentiable soft boundary mask
to extract content features (ck) viamask-weighted attention pooling,
which are then fused with the query via agated residual connection
to produce enriched queries (q˜k). -
Prediction Heads: These include boundary regression, validity classification, and the Mixture of Beta (MoB) uncertainty estimation model, which predicts parameters for a distribution capturing
the bimodal nature of uncertainty.
Training Strategy and Loss Functions
The training employs a two-phase strategy: a warmup phase training only span detection components followed by a joint phase activating all modules. The total loss function is complex, including:
-
Span Regression Loss (Lreg): Combines
L1 distance and Generalized IoU (GIoU) losses
over Hungarian-matched pairs. -
Validity Classification Loss (Lval): A binary cross-entropy loss to distinguish real spans from unused slots.
-
MoB Negative Log-Likelihood Loss (LMoB): Maximizes the likelihood of ground-truth uncertainty under the predicted mixture distribution, using parameters derived from the enriched query q˜k.
-
Contrastive Ranking Loss (Lrank): Preserves ordinal structure among span pairs by maximizing
the sign(u∗i − u∗j) · (ˆui − uˆj)
for stratified sampling of high- and low-uncertainty pairs.
Inference and Refinement
At inference time, SPANUQ uses a single forward pass followed by Uncertainty-Conditioned Iterative Refinement (UCIR). UCIR feeds Round 1 estimates back into the decoder for a second pass, where conditioned queries are refined using an MLP adapter. The final uncertainty estimate uˆk is obtained via a convex combination: uˆk = α · uˆ(2)k + (1 − α) · uˆ(1)k,
with α=0.7, adding less than 15% inference overhead.
Key Empirical Findings
Experiments on five LLM backbones show that SPANUQ consistently achieves the best span-level uncertainty quality (AUROC 0.908–0.944). A crucial observation is the span-to-sequence decomposability,
where the learned importance-weighted span composition achieves ρseq = 0.839, suggesting span-level estimation subsumes sequence-level as a special case.
Furthermore, SPANUQ’s DETR detector attains 0.
Improvements for AI systems
Here are specific improvements to AI systems based on the SPANUQ framework, and what those improved systems can achieve:
-
The development of a lightweight, end-to-end uncertainty quantification probe (SPANUQ) that operates directly on frozen LLM hidden states rather than expensive multi-sample inference or external knowledge retrieval.
-
The construction of the SPANUQ-BENCH benchmark, providing the first dataset with continuous soft labels derived from multi-sample claim verification, allowing for direct training of uncertainty estimation models.
-
The integration of a DETR-style span decoder coupled with a Mixture of Beta (MoB) uncertainty head to simultaneously detect semantically coherent text spans and estimate their uncertainty via bimodal distribution modeling.
-
The implementation of Uncertainty-Conditioned Iterative Refinement (UCIR), which allows the model to refine its initial span predictions by conditioning the second pass on the first pass's estimates, thereby correcting systematic errors in uncertainty estimation.
-
A framework capable of decomposing high-level sequence-level uncertainty into interpretable, granular span-level components (span-to-sequence decomposability), providing actionable insights into where a model is factually weak within a generated text.
These improvements enable the following capabilities:
-
The ability to perform
trust assessment
on generated text with high precision. Systems can now identify which specific sub-phrases (spans) in a long response carry the highest risk of hallucination, allowing users to focus their fact-checking efforts exactly where needed, rather than sifting through the entire output or relying on an unreliable single sequence score. -
The creation of self-refining LLMs that can improve their own outputs through uncertainty feedback loops. This allows models to iteratively refine uncertain spans identified in a first pass, leading to more coherent and factually grounded generations without needing massive computational overhead for multiple sampling runs during inference.
-
Enhanced diagnostic capabilities for LLM behavior. Researchers can analyze the internal hidden states of LLMs using SPANUQ to pinpoint which layers are most responsible for encoding factual uncertainty, guiding future model architectures toward better knowledge representation in critical areas (e.g., mid-to-late layers).
-
Robust deployment in high-stakes domains (legal, medical) by providing a calibrated confidence score that is granular enough to distinguish between a confidently correct term and an uncertain but potentially dangerous one, enabling human experts to make informed decisions based on the specific uncertainty of each claim.
-
Automated and efficient quality assurance pipelines. By using SPANUQ-BENCH labels, organizations can build automated systems that automatically flag generated content for manual review based on a quantified risk score derived from span-level uncertainty, significantly reducing the cost and time associated with human auditing.
Abstract
Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granularities: token-level scores lack semantic coherence, while sequence-level scores fail to localize errors. We formalize Span-Level Uncertainty Estimation (SLUE), a new task that targets the natural granularity for uncertainty: semantically coherent text spans, each conveying a single assessable unit of meaning. To address this task, we introduce SPANUQ, a lightweight (25M parameter) probe that distills the uncertainty knowledge from expensive multi-sample inference into a single forward pass over LLM hidden states. SPANUQ employs a DETR-style span decoder to simultaneously detect spans and estimate their uncertainty via a Mixture of Beta distribution, trained with a principled combination of Beta NLL regression and contrastive ranking objectives. We construct SPANUQ-BENCH, the first span-level uncertainty benchmark comprising 20K prompts, 293K annotated spans, and continuous soft labels derived from multi-sample claim verification. Experiments on five LLM backbones show that SPANUQ consistently achieves the best span-level uncertainty quality, outperforming the strongest probe baseline and all sampling-based methods while being 10 20x faster. Its DETR-based span detector attains 0.910 F1, surpassing the best heuristic by 39.4%, enabling precise error localization that sequence-level methods cannot provide. The same architecture ports to five LLMs spanning two model families, with one probe trained per backbone, and we additionally observe that sequence-level uncertainty is partially decomposable, suggesting that span-level estimation subsumes sequence-level as a special case. The project page is available damon-demon.github.io/SpanUQ.
Sources
- Language Models (Mostly) Know What They Know
- TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- To Believe or Not to Believe Your LLM
- LLM Uncertainty Quantification through Directional Entailment Graph and Claim Level Response Augmentation
- INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
- UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
- Fine-grained Hallucination Detection and Editing for Language Models
- CLUE: Concept-Level Uncertainty Estimation for Large Language Models
- HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering