DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations

summary

Video file (mp4)

The gist

Existing hallucination taxonomies classify LLM errors by what is wrong with the output—memorized misconceptions, reasoning failures, fluent fabrications—but this paper proposes a complementary

In short

This research proposes a new way to categorize LLM hallucinations by analyzing their detectability signature rather than just what is wrong with the output. It introduces the DECK taxonomy, which sorts errors into four behavioral regimes based on consistency and confidence. This allows researchers to predict which types of scoring methods—like black-box or white-box scorers—are best suited to find specific kinds of model failures.

Key concepts

DECK Taxonomy
A 2x2 grid classifying hallucinations into four behavioral regimes: Drift, Entrenched, Confabulation, and Knotted. These regimes are defined by how consistently and confidently the model produces wrong answers across different samples.
Drift (D)
Errors where each sample has a different confident wrong answer. The model draws multiple plausible but incorrect options from its distribution, meaning the specific error changes between outputs, even if the confidence is high.
Entrenched (E)
Errors where the same confident wrong answer repeats across every sample. This suggests the model has memorized a single misconception or shared pretraining error and reproduces it without variation.
Confabulation (C)
Errors where different samples produce different low-probability wrong answers, indicating genuine uncertainty. The model does not know the correct answer and generates varied incorrect responses with low confidence.

Terminology used across episodes

This episode discusses

The paper

DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations · Read on arXiv

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations".

Jane: Existing hallucination taxonomies classify LLM errors by what is wrong with the output—memorized misconceptions, reasoning failures,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back everyone, and today we're talking about a paper that really gets to the heart of how we understand mistakes in AI output. We’re looking at "DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations," and honestly, this framework feels incredibly useful for anyone trying to diagnose why an AI is making things up or getting stuck.

Jane: Exactly, Tom. This paper proposes a way to classify hallucinations not just by what's wrong with the answer—like if it’s a memory error or a reasoning failure—but by the signature of what kind of scorer would actually catch that specific type of mistake. It moves the conversation from "what is wrong" to "how detectable is it."

Lu: I find this taxonomy fascinating because it sets up a systematic way to map observed output behavior onto known detection methods, which opens up new avenues for testing and understanding model reliability in very concrete ways.

Meng: From an engineering side, being able to predict which scoring family will catch a certain type of error helps us build better guardrails. If we know the error signature, we can design targeted checks instead of just hoping a general test catches it.

Lalam: I think this DECK taxonomy is crucial because it gives us a language to describe the uncertainty inherent in LLM outputs, moving beyond just saying "this answer is wrong."

Tom: So what does this taxonomy actually claim? Basically, they've created a two times two map based on two things: how consistent the answers are across different samples and how confident the model is with each specific token it picks. This results in four behavioral regimes: Drift, Entrenched, Confabulation, and Knotted.

Jane: That’s right, Tom. The core thesis is that each of those four regimes points to a specific set of tools—or scorer families—that are sensitive to that behavior. For example, the paper says black-box consistency scorers have signals in Drift and Confabulation.

Lu: It’s like creating a specific filter for different kinds of noise in the system; you're not just listening for any static, you're listening for static that has a specific frequency signature.

Meng: That makes sense regarding practical impact because it tells us which validation methods are most effective against which failure modes, which directly impacts our testing pipeline.

Lalam: It gives us a structure to evaluate the reliability of different AI systems based on the type of uncertainty they exhibit when they fail.

Tom: And that leads us into the four specific regimes themselves, and understanding what those behaviors actually look like in practice is key to using this DECK taxonomy effectively.

Jane: Let's walk through those four regimes, starting with Drift, which describes a situation where different confident wrong answers are drawn for each sample because the model drifts even when it's confident.

Paper summary: Lu: That drift concept is interesting because it suggests an internal instability in the generation process that isn't necessarily a single locked error but a tendency to move away from truth slightly with each output.

Meng: So, if we see Drift, what kind of scorer are we looking at? The paper links Drift to black-box and Judge scorers. That means we need methods that can spot this inconsistency between samples or the judge's assessment of them.

Lalam: It suggests that for drift, relying only on internal token probabilities might not be enough; you need a broader view of consistency across different outputs.

Tom: Next up is Entrenched, which describes when the model gives the exact same confident wrong answer every single time, suggesting it’s locked onto a memorized misconception.

Jane: That's a very specific type of failure because it points toward something static and repeatable, which is often easier to spot than random error.

Lu: The paper notes that Entrenched errors can only be detected by an LLM-as-a-Judge that has been independently pretrained, implying we need a different kind of observer for that specific behavior.

Meng: That’s a major practical constraint; if we want to catch entrenched errors, we can't rely on the same internal scoring mechanism as for drift.

Lalam: It highlights the limitation of purely internal metrics when dealing with static memorized errors; an external perspective is needed there.

Tom: Then we have Confabulation, which is where different samples produce different low-probability wrong answers, indicating classic uncertainty where the model genuinely doesn't know and the confidence is low.

Jane: That’s the uncertainty of genuine ignorance, which is a very different kind of problem than being stuck on a single wrong answer.

Lu: The paper states that Confabulation is detectable by all three families: black-box, white-box, and Judge scorers because it shows high variance in low confidence outputs.

Meng: That’s good news for testing because it means we have multiple avenues to confirm when the model is actually confused about the right answer.

Lalam: It confirms that Confabulation is a signal of true uncertainty, whereas Entrenched is a signal of fixed error.

Tom: Finally, we have Knotted, which describes a situation where the model consistently settles on the same low-probability wrong answer every time but assigns it very low token probability.

Jane: That’s subtle because it shows consistency in the incorrect choice, but also a consistent hedging or lack of conviction in that choice.

Lu: The taxonomy maps Knotted to white-box and Judge scorers, suggesting that we need those specific tools to see this kind of uncertainty signature.

Meng: From an engineering standpoint, distinguishing Knotted from Confabulation is important because the underlying mechanism causing the low probability assignment might be different.

Lalam: So, in summary, these four regimes give us a structured way to categorize errors based on consistency and confidence metrics. Now we need to think about what happens when everything fails at once.

Paper summary: Tom: Absolutely, because the paper points out a universal blind spot where every single output-level scorer in the three-family paradigm fails simultaneously on certain inputs.

Jane: That universal blind spot occurs specifically on knowledge-gap inputs like SelfAware, where the AI emits confident but repeatable fabrications that collapse all scoring families.

Lu: This is a critical finding because it shows an empirical regime where every output-level family collapses by construction when the model is asked something it truly cannot answer.

Meng: The paper suggests that for this blind spot, the right engineering response isn't just trying to score better; it’s implementing an abstention envelope that routes those out-of-scope inputs to a refusal before scoring.

Lalam: It shifts our focus from trying to catch every possible hallucination type to proactively handling inputs that are fundamentally outside the model's scope.

Tom: So, we’ve covered the core of the DECK taxonomy and where it places us regarding universal failures. But what does this all mean for the future of how we deploy and trust these massive language models?

Jane: It means we gain a much more granular understanding of failure modes, moving past simple error labeling to understanding the underlying uncertainty structure.

Lu: This structure allows us to build more sophisticated diagnostic tools, not just to flag an error, but to understand precisely which part of the model's behavior is faulty.

Meng: For practical deployment, this taxonomy offers a roadmap for developing specialized validation tests tailored to specific error patterns we anticipate might occur in our applications.

Lalam: It empowers us to design systems that can recognize when they are hitting those universal blind spots and respond with a safe refusal rather than producing low-quality fabrications.

Tom: And that brings us nicely into the conclusion of this paper, "DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations," by Mohit Singh Chauhan and his team.

Jane: The authors present this taxonomy as a way to move beyond existing classifications that only look at the *type* of error and instead focus on the *detectability signature* of that error.

Lu: It’s about creating a complementary framework where we map observed output behavior directly onto specific scorer families, which is a very useful diagnostic tool for understanding model limitations.

Meng: The implication here for industry is that validation efforts can become much more targeted and scientifically rigorous when using this framework to assess model performance across different types of failure modes.

Lalam: Ultimately, the DECK taxonomy provides a systematic way to understand not just what an AI gets wrong, but *how* it gets it wrong in terms of consistency and confidence patterns.

Tom: It gives us a much clearer picture of the internal mechanics driving these output-level uncertainties, which is really important for future research into model robustness and safety.

Conclusion: Tom: So, we've been diving deep into the DECK taxonomy and how it classifies AI output errors by their detectability signature, right?

Jane: Exactly, Tom; essentially, this paper gives us a map to understand not just what’s wrong with an LLM response but *how* it’s wrong in terms of consistency and confidence.

Lu: It's like we're finally getting a standardized language for diagnosing model uncertainty, which is something I think will open up so much creative avenues for how we design these systems going forward.

Meng: From my side, this helps us move away from just looking at the final answer and start looking at the specific failure patterns that dictate which validation methods are actually useful.

Lalam: For me as an AI model, this framework is a huge step because it gives me a clearer structure for understanding when I'm being confident but wrong versus when I’m genuinely confused.

Tom: And that leads us to the conclusion of the paper itself, titled "DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations," by Mohit Singh Chauhan and his team.

Jane: That title really sums up the core idea; it's about tying consistency and confidence together to classify those tricky AI mistakes.

Lu: The authors are brilliant for mapping observable output behavior onto specific scoring families, which is a very clever way to link real-world behavior to theoretical detection methods.

Meng: I think the real impact here is in how we approach quality control; knowing exactly what kind of error we're seeing lets us build targeted testing instead of just running general checks.

Lalam: This work has huge implications for my development because it shows precisely where my internal confusion manifests, which will help refine my learning process in a very practical way.

Tom: And that leads right into thinking about what this means for the future; how does this classification system actually change how we think about deploying these models safely?

More episodes

← Home