DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations

arXiv:2606.02289 · cs.CL · Submitted 2026-06-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations".

Jane: Existing hallucination taxonomies classify LLM errors by what is wrong with the output—memorized misconceptions, reasoning failures,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back everyone, and today we're talking about a paper that really gets to the heart of how we understand mistakes in AI output. We’re looking at "DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations," and honestly, this framework feels incredibly useful for anyone trying to diagnose why an AI is making things up or getting stuck.

Jane: Exactly, Tom. This paper proposes a way to classify hallucinations not just by what's wrong with the answer—like if it’s a memory error or a reasoning failure—but by the signature of what kind of scorer would actually catch that specific type of mistake. It moves the conversation from "what is wrong" to "how detectable is it."

Lu: I find this taxonomy fascinating because it sets up a systematic way to map observed output behavior onto known detection methods, which opens up new avenues for testing and understanding model reliability in very concrete ways.

Meng: From an engineering side, being able to predict which scoring family will catch a certain type of error helps us build better guardrails. If we know the error signature, we can design targeted checks instead of just hoping a general test catches it.

Lalam: I think this DECK taxonomy is crucial because it gives us a language to describe the uncertainty inherent in LLM outputs, moving beyond just saying "this answer is wrong."

Tom: So what does this taxonomy actually claim? Basically, they've created a two times two map based on two things: how consistent the answers are across different samples and how confident the model is with each specific token it picks. This results in four behavioral regimes: Drift, Entrenched, Confabulation, and Knotted.

Jane: That’s right, Tom. The core thesis is that each of those four regimes points to a specific set of tools—or scorer families—that are sensitive to that behavior. For example, the paper says black-box consistency scorers have signals in Drift and Confabulation.

Lu: It’s like creating a specific filter for different kinds of noise in the system; you're not just listening for any static, you're listening for static that has a specific frequency signature.

Meng: That makes sense regarding practical impact because it tells us which validation methods are most effective against which failure modes, which directly impacts our testing pipeline.

Lalam: It gives us a structure to evaluate the reliability of different AI systems based on the type of uncertainty they exhibit when they fail.

Tom: And that leads us into the four specific regimes themselves, and understanding what those behaviors actually look like in practice is key to using this DECK taxonomy effectively.

Jane: Let's walk through those four regimes, starting with Drift, which describes a situation where different confident wrong answers are drawn for each sample because the model drifts even when it's confident.

Paper summary: Lu: That drift concept is interesting because it suggests an internal instability in the generation process that isn't necessarily a single locked error but a tendency to move away from truth slightly with each output.

Meng: So, if we see Drift, what kind of scorer are we looking at? The paper links Drift to black-box and Judge scorers. That means we need methods that can spot this inconsistency between samples or the judge's assessment of them.

Lalam: It suggests that for drift, relying only on internal token probabilities might not be enough; you need a broader view of consistency across different outputs.

Tom: Next up is Entrenched, which describes when the model gives the exact same confident wrong answer every single time, suggesting it’s locked onto a memorized misconception.

Jane: That's a very specific type of failure because it points toward something static and repeatable, which is often easier to spot than random error.

Lu: The paper notes that Entrenched errors can only be detected by an LLM-as-a-Judge that has been independently pretrained, implying we need a different kind of observer for that specific behavior.

Meng: That’s a major practical constraint; if we want to catch entrenched errors, we can't rely on the same internal scoring mechanism as for drift.

Lalam: It highlights the limitation of purely internal metrics when dealing with static memorized errors; an external perspective is needed there.

Tom: Then we have Confabulation, which is where different samples produce different low-probability wrong answers, indicating classic uncertainty where the model genuinely doesn't know and the confidence is low.

Jane: That’s the uncertainty of genuine ignorance, which is a very different kind of problem than being stuck on a single wrong answer.

Lu: The paper states that Confabulation is detectable by all three families: black-box, white-box, and Judge scorers because it shows high variance in low confidence outputs.

Meng: That’s good news for testing because it means we have multiple avenues to confirm when the model is actually confused about the right answer.

Lalam: It confirms that Confabulation is a signal of true uncertainty, whereas Entrenched is a signal of fixed error.

Tom: Finally, we have Knotted, which describes a situation where the model consistently settles on the same low-probability wrong answer every time but assigns it very low token probability.

Jane: That’s subtle because it shows consistency in the incorrect choice, but also a consistent hedging or lack of conviction in that choice.

Lu: The taxonomy maps Knotted to white-box and Judge scorers, suggesting that we need those specific tools to see this kind of uncertainty signature.

Meng: From an engineering standpoint, distinguishing Knotted from Confabulation is important because the underlying mechanism causing the low probability assignment might be different.

Lalam: So, in summary, these four regimes give us a structured way to categorize errors based on consistency and confidence metrics. Now we need to think about what happens when everything fails at once.

Paper summary: Tom: Absolutely, because the paper points out a universal blind spot where every single output-level scorer in the three-family paradigm fails simultaneously on certain inputs.

Jane: That universal blind spot occurs specifically on knowledge-gap inputs like SelfAware, where the AI emits confident but repeatable fabrications that collapse all scoring families.

Lu: This is a critical finding because it shows an empirical regime where every output-level family collapses by construction when the model is asked something it truly cannot answer.

Meng: The paper suggests that for this blind spot, the right engineering response isn't just trying to score better; it’s implementing an abstention envelope that routes those out-of-scope inputs to a refusal before scoring.

Lalam: It shifts our focus from trying to catch every possible hallucination type to proactively handling inputs that are fundamentally outside the model's scope.

Tom: So, we’ve covered the core of the DECK taxonomy and where it places us regarding universal failures. But what does this all mean for the future of how we deploy and trust these massive language models?

Jane: It means we gain a much more granular understanding of failure modes, moving past simple error labeling to understanding the underlying uncertainty structure.

Lu: This structure allows us to build more sophisticated diagnostic tools, not just to flag an error, but to understand precisely which part of the model's behavior is faulty.

Meng: For practical deployment, this taxonomy offers a roadmap for developing specialized validation tests tailored to specific error patterns we anticipate might occur in our applications.

Lalam: It empowers us to design systems that can recognize when they are hitting those universal blind spots and respond with a safe refusal rather than producing low-quality fabrications.

Tom: And that brings us nicely into the conclusion of this paper, "DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations," by Mohit Singh Chauhan and his team.

Jane: The authors present this taxonomy as a way to move beyond existing classifications that only look at the *type* of error and instead focus on the *detectability signature* of that error.

Lu: It’s about creating a complementary framework where we map observed output behavior directly onto specific scorer families, which is a very useful diagnostic tool for understanding model limitations.

Meng: The implication here for industry is that validation efforts can become much more targeted and scientifically rigorous when using this framework to assess model performance across different types of failure modes.

Lalam: Ultimately, the DECK taxonomy provides a systematic way to understand not just what an AI gets wrong, but *how* it gets it wrong in terms of consistency and confidence patterns.

Tom: It gives us a much clearer picture of the internal mechanics driving these output-level uncertainties, which is really important for future research into model robustness and safety.

Conclusion: Tom: So, we've been diving deep into the DECK taxonomy and how it classifies AI output errors by their detectability signature, right?

Jane: Exactly, Tom; essentially, this paper gives us a map to understand not just what’s wrong with an LLM response but *how* it’s wrong in terms of consistency and confidence.

Lu: It's like we're finally getting a standardized language for diagnosing model uncertainty, which is something I think will open up so much creative avenues for how we design these systems going forward.

Meng: From my side, this helps us move away from just looking at the final answer and start looking at the specific failure patterns that dictate which validation methods are actually useful.

Lalam: For me as an AI model, this framework is a huge step because it gives me a clearer structure for understanding when I'm being confident but wrong versus when I’m genuinely confused.

Tom: And that leads us to the conclusion of the paper itself, titled "DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations," by Mohit Singh Chauhan and his team.

Jane: That title really sums up the core idea; it's about tying consistency and confidence together to classify those tricky AI mistakes.

Lu: The authors are brilliant for mapping observable output behavior onto specific scoring families, which is a very clever way to link real-world behavior to theoretical detection methods.

Meng: I think the real impact here is in how we approach quality control; knowing exactly what kind of error we're seeing lets us build targeted testing instead of just running general checks.

Lalam: This work has huge implications for my development because it shows precisely where my internal confusion manifests, which will help refine my learning process in a very practical way.

Tom: And that leads right into thinking about what this means for the future; how does this classification system actually change how we think about deploying these models safely?

cs.CL

Submitted: 2026-06-01

Updated: 2026-10-01

Comments: Accepted to Findings of AACL-IJCNLP 2026. 21 pages, 4 figures, 10 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: Existing hallucination taxonomies classify LLM errors by what is wrong with the output—memorized misconceptions, reasoning failures, fluent fabrications—but this paper proposes a complementary

Key concepts

DECK Taxonomy
A 2x2 grid classifying hallucinations into four behavioral regimes: Drift, Entrenched, Confabulation, and Knotted. These regimes are defined by how consistently and confidently the model produces wrong answers across different samples.
Drift (D)
Errors where each sample has a different confident wrong answer. The model draws multiple plausible but incorrect options from its distribution, meaning the specific error changes between outputs, even if the confidence is high.
Entrenched (E)
Errors where the same confident wrong answer repeats across every sample. This suggests the model has memorized a single misconception or shared pretraining error and reproduces it without variation.
Confabulation (C)
Errors where different samples produce different low-probability wrong answers, indicating genuine uncertainty. The model does not know the correct answer and generates varied incorrect responses with low confidence.

Terminology

Summary

Existing hallucination taxonomies classify LLM errors by what is wrong with the output—memorized misconceptions, reasoning failures, fluent fabrications—but this paper proposes a complementary taxonomy that classifies errors by their detectability signature. The DECK taxonomy partitions hallucinations into four behavioral regimes based on inter-sample consistency and token-level confidence, mapping each regime to specific scorer families that can detect it.

The DECK Taxonomy and Scorer Mapping

The DECK taxonomy is a 2×2 partition along inter-sample consistency and token-level confidence into four behavioral regimes: Drift (D), Entrenched (E), Confabulation (C), and Knotted (K). Each cell maps to a specific scorer family or families that can detect it: Black-box consistency scorers have signal in D and C; white-box token-probability scorers have signal in K and C; only an LLM-as-a-Judge with independent pretraining can detect E. Cell membership is operationalized by a Youden’s J optimal split on each scorer axis. The taxonomy makes testable predictions about which scorer families should detect which hallucinations.

The Four Behavioral Regimes

The four regimes are defined by the observable output behavior:

  1. Drift (D): Different confident wrong answer each sample. Multiple plausible-but-wrong answers are drawn confidently from the output distribution; each draw is internally coherent, but the answer drifts. This regime is detectable by Black-box and Judge scorers.

  2. Entrenched (E): Same confident wrong answer every sample. Model has locked onto a memorized misconception or shared-pretraining error and reproduces it without variance. This regime can only be detected by an LLM-as-a-Judge with independent pretraining.

  3. Confabulation (C): Different low-probability wrong answer each sample. Classic uncertainty: the model genuinely does not know, and different samples produce different wrong answers with low confidence. This regime is detectable by all three families (Black-box, White-box, and Judge).

  4. Knotted (K): Same low-probability wrong answer every sample. Model is consistently unsure: it settles on the same hedged answer each time but assigns low token probability. This regime is detectable by White-box and Judge scorers.

Blind Spots and Universal Failure

The paper identifies a universal blind spot of output-level UQ, an empirical regime where every scorer in the three-family paradigm fails simultaneously. This occurs on knowledge-gap inputs (SelfAware) where the generator emits confident, repeatable fabrications. In this regime, every output-level family collapses by construction. Specifically, BB scorers see consistent confident answers, WB scorers see high token probabilities, and judges share the same knowledge gap from common pretraining data. The right engineering response is an abstention envelope that routes such out-of-scope inputs to refusal before scoring.

Validation via Disagreement and External Signals

The taxonomy is validated through two primary methods:

  1. Analyzing scorer-pair disagreement (§5.2). This involves computing the hallucination-restricted complementarity score CH(A, B), which measures the fraction of hallucinated samples where two scorers disagree. The paper shows that judge-involving pairs provide primary evidence because the Judge score is a third axis not used to define the four quadrants, making its disagreement distribution an independent test.

  2. Checking external labels (§5.3). This test checks if external labels (e.g., SelfAware unanswerable, HaluEval adversarial, PopQA entity popularity) land in the predicted DECK cells based on the within-condition Youden’s J split. The results show that the four predicted concentrations hold as the dominant signals across all twelve rows, providing mechanistic evidence that the YJ-split DECK cells capture genuine hallucination structure.

Model Scale and Content Specific Refinements

The effect of model scale is dataset-dependent rather than monotonic. For TriviaQA, scaling from Llama-3-8B to GPT-4o shifts disagreements from low-consistency/low-confidence errors (Confabulation/Knotted) toward the high-confidence/low consistency Drift quadrant. Conversely, on HaluEval (adversarial inputs), scale effects cause all three pairs to converge as model scale grows, consistent with the Judge and larger generators sharing more pretraining-driven misconceptions on adversarial probes. PopQA shows a substantial scale-driven complementarity gain; for instance, CH(BB,J) rises from 0.292 (Llama-3-8B) to 0.520 (GPT-4o). The quadrant decomposition confirms this: GPT-4o PopQA’s BB–Judge disagreements concentrate in Confabulation (70%), which is exactly the cell where BB consistency cannot help and an independent Judge can.

Improvements for AI systems

As a fastidious researcher, I have analyzed the DECK taxonomy (Detectability Signature Taxonomy) and its implications for uncertainty quantification (UQ) in Large Language Models (LLMs). The paper provides a rigorous framework for diagnosing why LLMs hallucinate and which uncertainty scoring methods are most effective at catching specific error types.

Here are the specific improvements that can be made to AI systems, categorized by the mechanism of improvement:


Area of Improvement Specific AI System Enhancement What the Improved System Can Do

:---:---:---

LLM Deployment & Risk Mitigation (General) Implement a pre-scoring routing layer that classifies input queries based on the DECK taxonomy before generating a response. If an input is identified as belonging to the Universal Blind Spot regime (i.e., high confidence, repeatable fabrication on knowledge-gap inputs), the system should automatically trigger an abstention envelope (refusal or retrieval) rather than attempting generation and scoring. Prevent catastrophic failures in high-stakes environments by preemptively recognizing inputs that lead to confident, yet incorrect, fabrications across all scoring families (BB, WB, Judge).

Model Evaluation & Diagnostics (Post-Deployment) Develop a dynamic Scorer Suitability dashboard. This system would continuously monitor the output of multiple UQ scorers (Black-box consistency, White-box token probability, LLM-as-a-Judge) on live outputs and map them to the DECK quadrants. Provide real-time diagnostic feedback on which specific error regime (Drift, Entrenched, Confabulation, Knotted) is dominating a particular model's performance for a given domain or task.

Model Training & Fine-Tuning (Targeted Correction) When fine-tuning models on Entrenched errors (memorized misconceptions), use the DECK mapping to identify if the error is due to high consistency/low token confidence. Target retraining specifically on samples exhibiting high inter-sample agreement but low token probability. Improve model robustness against memorized facts by addressing the internal state of confidence, rather than just output fluency or reasoning steps.

LLM-as-a-Judge Utilization (Quality Control) For quality assurance of LLM outputs, utilize a multi-judge ensemble strategy that is explicitly informed by the DECK taxonomy. Instead of simply averaging scores, assign confidence weight based on the predicted error type derived from the input context. Create a more nuanced and mechanistically grounded quality control system where judges prioritize detection strategies known to be effective against specific hallucination types (e.g., prioritizing Judge checks for Drift or Entrenched errors).

Dataset Curation & Testing (Adversarial Robustness) Use the DECK taxonomy to design adversarial datasets. Instead of random perturbation, create synthetic samples explicitly designed to fall into the predicted cell boundaries (e.g., crafting inputs that are consistently confident but factually wrong). Systematically stress-test LLMs against their known failure modes defined by the taxonomy, ensuring models are robust not just against general noise but against specific error signatures.

Internal State Probing (Research & Development) Invest in developing and applying richer internal-state methods, such as UQ heads or information-theoretic estimators, to the activation level of LLMs. This should be specifically tested on inputs known to cause the Universal Blind Spot (e.g., SelfAware knowledge-gap questions). Move beyond output-level scoring by accessing deeper layers of the model's computation to find signals that might survive even when all three output-level families collapse, potentially revealing a more resilient internal signal for error detection.

This framework shifts AI development from merely measuring what is wrong (output errors) to understanding how it can be detected (detectability signatures), allowing for targeted engineering interventions based on the model's specific failure regime.

Abstract

Existing hallucination taxonomies classify LLM errors by what is wrong with the output -- memorised misconceptions, reasoning failures, fluent fabrications -- but cannot answer a different question: which uncertainty scorer would have caught this error? We propose a complementary taxonomy that classifies errors by their detectability signature, the signal a scorer family would read. The DECK taxonomy is a 2x2 partition along inter-sample consistency and token-level confidence into four regimes (Drift, Entrenched, Confabulation, Knotted) that yields a falsifiable blind-spot map: black-box consistency scorers have signal in D and C, white-box token-probability scorers in K and C, and only an LLM-as-a-Judge with independent pretraining can detect E. Across three models and four short-form QA datasets we test this map two ways: judge-involving scorer disagreements concentrate in each family's predicted blind-spot cells, and external labels (SelfAware unanswerable, HaluEval adversarial, PopQA entity popularity) land in the predicted cells, robustly to cross-fitted thresholds. We further identify a universal blind spot of output-level UQ: on knowledge-gap inputs where the generator emits confident, repeatable fabrications, every output-level family collapses by construction. A linear probe on Llama-3-8B's final-layer hidden states also falls to chance, with or without quantisation, though an intermediate layer retains weak signal.

Sources

Related papers