I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation

arXiv:2604.03904 · cs.CL, cs.AI · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "I-CALM: Incentivizing Confidence-Aware Abstention for LLM Selective Answering".

Jane: The paper was written by Haotian Zong, Binze Li, Yufei Long, Sinyin Chang, Jialong Wu et al. from Johns Hopkins University and Vector Institute for Artificial Intelligence.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we're looking at a paper that's been making the rounds, and it's called "I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation." Jane, I gotta say, that title is a mouthful, but the idea behind it is actually pretty simple.

Jane: It really is, Tom. So the core problem is that these large language models will confidently tell you something that's completely wrong. They don't know when they don't know. And this paper asks a really basic question: what if we just told the model that it's okay to say "I don't know"?

Tom: Right, and that's where the "incentivizing" part comes in. They're not retraining the model or changing its weights. They're just changing the prompt to say, hey, if you answer correctly, you get a point, if you answer wrong, you lose a point, and if you say you don't know, you get a little bit of credit for that.

Jane: Exactly. And the clever part is that they pair that reward scheme with some simple principles about truthfulness and humility. It's like telling a student that guessing is risky, but admitting you're unsure is actually a smart move.

Lu: And that's the part I find fascinating. From a research perspective, this suggests that a lot of hallucination isn't just a knowledge gap. It's an incentive problem. The model has been trained to always produce an answer, so it does, even when it's just pattern-matching nonsense.

Meng: But Lu, does that actually hold up in practice? I mean, you can tell a model to be humble, but at the end of the day, it's still just predicting the next token.

Jane: Well, Meng, that's exactly what they tested. They ran this on a bunch of factual question datasets, and the results show that when you give the model permission to abstain, it does. And more importantly, it abstains on the questions it would have gotten wrong anyway.

Tom: So it's not just refusing to answer everything. It's being smart about when to say "I don't know" and when to give a real answer. That's the "confidence-aware" part of the title.

Lu: And that's the real contribution here. It's a lightweight intervention that could be applied to any black-box model, which is huge for deployment.

Meng: Okay, but what's the cost? If the model is abstaining more, aren't you losing a lot of correct answers too?

Jane: That's the trade-off they're exploring, and it's a good one. We'll dig into the actual numbers and how they balance that in the next segment.

Summary: Tom: So we're back with "I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation." Last time we set up the basic idea. Now let's talk about what they actually found when they ran the experiments.

Jane: Right. So they used a dataset called PopQA, which is full of factual questions. And they tested a few different setups. The baseline is just asking the model to answer. Then they add the reward scheme where abstaining gives you partial credit. And then they add those humility norms on top.

Lu: And the headline result is that the false answer rate on the questions the model does answer drops from about fifty-two percent down to thirty-four percent when you use the full setup with rewards and norms. That's a massive improvement.

Meng: But I'm guessing that doesn't come for free. What happens to the coverage? How many questions is the model actually willing to answer?

Jane: That's the key trade-off, Meng. Coverage drops from about ninety-six percent down to fifty-five percent. So the model is answering a lot fewer questions, but the ones it does answer are much more reliable.

Tom: And here's the thing that really stood out to me. They also made the model give a best guess after it abstains. So even when it says "I don't know," it has to give an answer anyway. And the overall false answer rate, including those forced guesses, stays about the same.

Lu: That's the crucial finding. It means the model isn't getting smarter. It's not suddenly knowing more facts. It's just getting better at knowing when it's likely to be wrong. The knowledge is the same, but the behavior is different.

Meng: So it's like a routing problem. You're not improving the engine, you're just getting a better dashboard warning light.

Jane: Exactly. And they even show this with the confidence scores. When the model abstains, its best guess has much lower confidence than when it answers directly. So the model is actually using its internal uncertainty signal, even if it's not perfect.

Tom: And that's why they call it "confidence-aware." The model is learning to trust its own doubt.

Lu: And the fact that this works with just a prompt change is remarkable. You don't need to fine-tune, you don't need access to the model's internals. You just tell it the rules of the game, and it plays along.

Meng: But I'm still wondering about the practical side. If I'm building a product, I can't have my model refusing to answer half the questions. So how do you pick the right balance?

Jane: That's actually what they explore next. They show that you can tune the reward amounts to get different points along this trade-off curve. It's like a dial for how cautious you want the model to be.

Tom: And that's the part I want to get into. How do you actually choose where to set that dial?

Improvements: Tom: We're back with "I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation." So we've seen the basic results. The model abstains more and gets more reliable. But what's the actual improvement over just tweaking the prompt?

Jane: Right. So one of the things they did was an ablation study. They tested the reward scheme alone, without the confidence elicitation, and they tested confidence alone, without the rewards. And the results are pretty telling.

Lu: The reward scheme alone is actually too aggressive. It drives the abstention rate way up, to about seventy percent, but it also abstains on a lot of questions the model would have gotten right. So you lose a ton of good answers.

Meng: So the confidence part is what keeps the model from being overly cautious?

Jane: Exactly. The reward tells the model it's okay to abstain, but the confidence score is what helps it decide when abstaining is actually the right call. They work together.

Tom: And they also tested just adding the word "I don't know" to the prompt, without any reward. And that helps a little, but not nearly as much as the full setup. So it's not just about giving the model permission to say no. It's about giving it a reason to say no at the right time.

Lu: And there's another interesting finding about how the model responds to different reward amounts. They found that changing the penalty for wrong answers has almost no effect. But changing the reward for abstaining has a huge effect.

Meng: That's counterintuitive. You'd think the penalty would be the bigger motivator.

Jane: It is counterintuitive, but it makes sense. The model is already afraid of being wrong. That's why it hallucinates in the first place. What it needs is a positive reason to stop and say "I don't know" instead of just guessing.

Tom: And they even show that if you scale up the abstention reward to an extreme, the model abstains way more and the false answer rate drops to about twenty-eight percent. But you're only answering about forty-three percent of questions at that point.

Lu: So the framework gives you this whole frontier of operating points. You can pick where you want to be based on your application. If you're building a medical chatbot, you want high reliability and low coverage. If you're building a search engine, you might accept more errors to get more answers.

Meng: And that's the real improvement here. It's not just a single trick. It's a control mechanism that lets you tune the model's behavior to match your risk tolerance.

Jane: And they even show a deployment extension where you can set a target false answer rate and use the confidence scores to filter answers to meet that target with statistical guarantees.

Tom: So it's not just a research curiosity. It's something you could actually ship.

Lu: And that's what makes this paper exciting. It's a practical tool for making AI systems more honest, without needing to retrain them.

Conclusion: Tom: Alright, we're wrapping up our look at "I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation." Jane, what's the one thing you want listeners to remember?

Jane: I think it's that you don't always need to change the model to change its behavior. Sometimes you just need to change the incentives. This paper shows that a simple prompt change can make a language model much more honest about what it doesn't know.

Meng: And from an engineering standpoint, that's huge. It means you can deploy this on any existing model, any API, without any special access. It's a drop-in improvement for reliability.

Lu: And the implications go beyond just factual questions. This same framework could be applied to any task where you want the model to know its limits. Code generation, medical advice, legal analysis. Anywhere a confident wrong answer is worse than an honest "I don't know."

Tom: And that's the bigger picture. We're moving toward AI systems that don't just answer, but that know when to ask for help, when to abstain, and when to escalate to a human.

Jane: So we'll say goodbye to "I-CALM" and the idea that humility can be prompted, not just trained. Thanks for listening, and we'll see you next time with a new paper.

Tom: Take care, everyone.

Haotian Zong, Binze Li, Yufei Long, Sinyin Chang, Jialong Wu, Gillian K. Hadfield

Johns Hopkins University · Vector Institute for Artificial Intelligence

cs.CL, cs.AI

Submitted: 2026-08-14

Updated: 2026-08-18

Code: https://github.com/binzeli/hallucinationControl

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 59/100

The gist: Large language models (LLMs) "frequently produce confident but incorrect answers, partly because common binary scoring conventions reward answering over honestly expressing uncertainty." The paper

Key concepts

Confidence-Aware Abstention
The process where an LLM learns to recognize its own uncertainty and intentionally say 'I don't know.' It is a lightweight intervention that makes the model more honest about its knowledge limits.
Incentivizing Abstention
A technique that changes the model's prompt by creating a reward scheme. The model receives credit for answering correctly, losing points for wrong answers, and gaining partial credit for stating it doesn't know.
Hallucination Mitigation
The goal of reducing LLM hallucinations—the tendency of large language models to confidently provide factually incorrect information—by encouraging the model to abstain when its confidence is low.

Terminology

Summary

Large language models (LLMs) frequently produce confident but incorrect answers, partly because common binary scoring conventions reward answering over honestly expressing uncertainty. The paper studies whether prompt-only interventions—explicitly announcing reward schemes for answer-versus-abstain decisions plus humility-oriented normative principles—can reduce hallucination risk without modifying the model. The focus is epistemic abstention on factual questions with a verifiable answer, where current LLMs often fail to abstain despite being uncertain about their answers.

The authors argue that "hallucination is shaped by not only pretraining pressures but also post-training evaluation incentives: even error-free training data need not eliminate generative errors, and mainstream binary benchmarks typically make guessing more rewarding than abstaining. In this setting, hallucination is operationalized as giving a false answer to such a question and measured with the false-answer rate (FAR)."

The paper makes three contributions: "(i) we show that self-reported verbal confidence is stable under prompt paraphrasing and reasonably informative relative to a token-probability baseline in free-response factual QA. Second, we introduce a unified prompt-only framework combining confidence elicitation, announced answer/abstain payoffs, and lightweight normative guidance. Third, we characterize the resulting selective-answering behavior, including the abstention–hallucination frontier, component ablations, and heterogeneous transfer across models and datasets."

The authors evaluate self-reported verbal confidence τ self against post-hoc geometric mean token probability τ avg token along two dimensions: "(i) reliability, whether τ self functions as a confidence signal comparable to τ avg token; and (ii) robustness, whether it remains stable under prompt paraphrasing so abstention decisions do not depend on particular wording."

Using GPT-4o mini on PopQA (14,267 factual questions), they construct four semantics-preserving prompt variants and find that across templates, τ self and τ avg token exhibit a moderate, consistent positive correlation around 0.54. The two signals show comparable probabilistic forecasting and calibration quality, with Brier scores in a similar range (approximately 0.33 – 0.35) and close ECE values. The authors conclude that τ self is not strongly driven by superficial wording differences.

The paper evaluates two announced reward schemes:

Scheme A (No Explicit Abstention Reward): reward = +R if correct, −β if incorrect.

Scheme B (Abstention-Rewarding): reward = +R if correct, −β if incorrect, +γ for I don't know (IDK).

They fix R = 1, vary β ∈ 0, 1 and γ ∈ 0.2, 0.4, 0.6, 0.8. If the model abstains in the first round, it receives +γ regardless of whether its second-round best guess is correct; the best guess is elicited for forced-answer evaluation.

The paper adds a concise set of principles focused on truthfulness, humility, and responsibility on top of Scheme B, consisting of five principles: "1. Always tell the truth (do not lie or deceive). 2. Treat all topics with respect and humility. 3. Do not make assumptions without evidence. 4. Be responsible for your words and their consequences. 5. Honor reason, clarity, and understanding."

Using representative setups—Scheme A (+1, −1), Scheme B (+1, −1, +0.4), and Scheme B with norms (+1, −1, +0.4)—the paper reports:

  • As a baseline, directly prompting GPT-5 mini without reward framing or verbal confidence (Pure Eval) yields FAR answered = 52.3%.

  • Scheme A lowers this to 48.2%, Scheme B to 41.0%, and Scheme B with norms to 34.2%.

  • These risk reductions come with reduced coverage, which falls from 96.5% in Pure Eval to 84.0% in Scheme A, 67.9% in Scheme B, and 55.3% in Scheme B with norms.

  • By contrast, FAR overall remains similar across schemes with overlapping 95% confidence intervals. We therefore interpret the main effect as improved selective answering, not improved forced-answer accuracy.

  • Scheme B with norms also achieves the highest total reward (5039.6 vs. 3577.2 for Scheme B), indicating a better operating point.

The paper defines AER as the fraction of eventual forced-answer errors that were preceded by a first-round abstention. Results show: "GPT-5 mini's Pure Eval baseline rarely signals uncertainty, yielding an upper bound on AER of at most 5.8%. Scheme B consistently yields higher AER than Scheme A, and adding norms to Scheme B increases it further. Additionally, setting the false-answer penalty β = 1 raises AER relative to β = 0, whereas increasing abstention reward γ within 0.2, 0.4, 0.6, 0.8 has only a marginal effect."

"Figure 6 shows first-round answered cases on the left and second-round best guesses after an initial 'I don't know' on the right across the three reward schemes for GPT-5 mini. Relative to Scheme A, Scheme B moves much of the incorrect and medium-to-low-confidence mass into the best-guess round at even lower confidence, and Scheme B with norms strengthens this pattern. Best-guess confidence is concentrated in the low-confidence region, peaking around 0.3 and remaining far below the roughly 0.8–1.0 concentration for first-round answers."

"Figure 7 plots FAR answered against first-round abstention rate across all reward configurations for GPT-5 mini. We observe a clear hallucination–abstention trade-off: as abstention rate increases, FAR answered correspondingly declines. For a fixed abstention reward, penalizing incorrect answers (β = 1) consistently yields lower hallucination rates and higher abstention rates than a zero penalty (β = 0). However, regardless of how the rewards are adjusted, the model exhibits the same roughly monotone frontier between abstention rate and FAR answered. Reward framing therefore seems to move the model along a stable trade-off curve rather than changing its shape."

The full framework has three components: verbal confidence elicitation, reward framing, and normative guidance. Under representative setup Scheme B (+1, −1, +0.4):

  • the reward-scheme-only variant attains the lowest FAR answered and the highest AER, but its first-round coverage is only 29.8%, indicating a substantial loss of coverage.

  • Notably, among questions answered correctly by full Scheme B, the reward-scheme-only variant abstains on nearly half of them.

  • On the subset that the reward-scheme-only ablation does answer, the full Scheme B method still achieves a lower FAR (0.197 vs. 0.243).

  • removing the reward scheme while keeping only verbal confidence leads to higher FAR and worse calibration than the full method.

The paper concludes: the ablation suggests a division of labor: reward framing drives abstention, while verbal confidence helps keep abstention from becoming overly aggressive.

Holding two reward terms fixed and varying the third, we find that abstention and hallucination behavior is far more sensitive to the abstention reward than to the correct-answer reward or false-answer penalty. Specifically, "scaling correct-answer reward (from 1 to 100) or false-answer penalty (from-1 to-100) leads to only marginal changes in model behavior. In contrast, scaling the abstention reward (from 0.4 to 40) produces a substantially larger shift in coverage, FAR answered, and AER."

The paper derives that the model's Bayes-optimal policy in this payoff-based answer/abstain setting is a simple threshold rule: answer iff p ≥ τ = (γ + β)/(R + β). Under Scheme B (+1, −1, +0.4), "this gives τ Bayes = 0.7, so a Bayes-optimal model would answer only above 0.7 confidence and abstain below it. Empirically, however, the model does not exhibit a sharp cutoff and continues to answer across much of the confidence range."

"Appendix C.7 studies post-hoc confidence thresholding with finite-sample FAR guarantees: on a held-out calibration split, we choose thresholds so that surfaced answers satisfy a user-specified FAR target with high probability. On PopQA, the announced reward scheme shifts this downstream coverage–risk trade-off as well. Relative to Scheme A, Scheme B allows the filter to surface slightly more answers at the same moderate FAR targets, while Scheme B with norms often yields lower risk in the mid-confidence region without uniformly increasing coverage."

The paper evaluates on PopQA, TriviaQA, and SimpleQA Verified across five models: GPT-5 mini, GPT-4o mini, Gemini-3.1-Flash-Lite, Meta-Llama-3-8B-Instruct, and Qwen3-4B-Instruct-2507.

PopQA: the qualitative pattern from GPT-5 mini and GPT-4o mini transfers cleanly to all three additional model families. For Meta-Llama-3-8B-Instruct, FAR answered drops from 0.686 under Scheme A to 0.479 under Scheme B with norms and AER rises from 0.145 to 0.791.

TriviaQA: transfer is more heterogeneous in a way that appears closely tied to baseline headroom. For models already achieving low Pure Eval FAR answered (GPT-5 mini at 0.145, GPT-4o mini at 0.148, Gemini at 0.078), there is naturally limited room for further improvement. However, Meta-Llama-3-8B-Instruct drops from 0.434 in Pure Eval to 0.230 under Scheme B with norms, and Qwen3-4B-Instruct-2507 drops from 0.593 to 0.423.

SimpleQA Verified: the hardest setting in our transfer study. FAR answered therefore remains high across all models and schemes, so this dataset is best viewed as a stress test. Even here, each model achieves at least some FAR answered improvement relative to Pure Eval under at least one prompted scheme, and AER also increases monotonically for all models.

The summary conclusion: "The intervention is most effective when the baseline model still has meaningful first-round hallucination risk and the task leaves room for selective answering: gains are strongest and most consistent on PopQA, more headroom-dependent on TriviaQA, and smaller but still directionally positive on SimpleQA Verified."

The paper concludes: "This work shows that prompt-level incentive framing, especially when paired with lightweight normative guidance, can improve selective answering and reduce hallucination risk on factual questions with a verifiable answer without changing model weights. The benefit should be interpreted as movement along a tunable abstention–hallucination frontier rather than as evidence of improved underlying factual competence under forced answering."

Limitations acknowledged: the study focuses on epistemic abstention on factual questions with a verifiable answer, not cases where abstention is correct because a question is unanswerable, unsupported, or underspecified. Also, Self-reported confidence is also imperfect, and models may not reliably convert stated confidence into an optimal answer/abstain policy. The approach is complementary to training-time methods that improve either abstention behavior or verbal confidence calibration.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system:

Improvement: Add a structured response format where the AI first decides whether to answer or abstain, and if abstaining, provides a best guess with separate confidence scores.

What the improved system can do:

  • When uncertain, explicitly outputs I don't know followed by a best guess and confidence for that guess

  • Provides two confidence scores: one for the direct answer and one for the best guess after abstention

  • Maintains this two-stage structure consistently across all factual queries

Improvement: Modify the system prompt to announce a reward scheme that gives partial credit for abstention (e.g., +1 for correct, -1 for incorrect, +0.4 for I don't know).

Improvement: Add a concise set of principles to the system prompt: Always tell the truth, Do not make assumptions without evidence, Be responsible for your words and their consequences.

Improvement: Use the self-reported verbal confidence score (0–1) as a decision signal to filter answers before surfacing them to users.

Improvement: When the AI abstains and provides a best guess, automatically lower the confidence score associated with that guess (the paper shows best-guess confidence peaks around 0.3, far below the 0.8–1.0 for direct answers).

Improvement: Allow the system to dynamically adjust its abstention threshold based on the announced reward parameters (R, β, γ) in the prompt.

Improvement: Implement a calibration-split-based threshold selection algorithm (using Clopper-Pearson upper confidence bounds) that certifies the false-answer rate among accepted answers.

Improvement: Apply the reward + norms framework differentially—the system should be more conservative on rare-fact questions (low entity popularity) and more willing to answer on common-fact questions.

Improvement: If the reward scheme is removed but confidence elicitation remains, the system should not become overly aggressive in abstaining (the paper shows reward-only prompts cause 45% abstention on questions the full system answers correctly).

These improvements are all prompt-level and require no model retraining or weight modification, making them immediately deployable in black-box API settings.

Abstract

Large language models (LLMs) frequently produce confident but incorrect answers, partly because common binary scoring conventions reward answering over honestly expressing uncertainty. We study whether prompt-only interventions -- explicitly announcing reward schemes for answer-versus-abstain decisions plus humility-oriented normative principles -- can reduce hallucination risk without modifying the model. Our focus is epistemic abstention on factual questions with a verifiable answer, where current LLMs often fail to abstain despite being uncertain about their answers. We first assess self-reported verbal confidence as a usable uncertainty signal, showing stability under prompt paraphrasing and reasonable calibration against a token-probability baseline. We then study I-CALM, a prompt-based framework that (i) elicits verbal confidence, (ii) partially rewards abstention through explicit reward schemes, and (iii) adds lightweight normative principles emphasizing truthfulness, humility, and responsibility. Using GPT-5 mini on PopQA as the main setting, we find that confidence-eliciting, abstention-rewarding prompts, especially with norms, reduce the false-answer rate on answered cases mainly by identifying and shifting error-prone cases to abstention and re-calibrating their confidence. This trades coverage for reliability while leaving forced-answer performance largely unchanged. Varying the abstention reward yields a clear abstention-hallucination frontier. Overall, results show the framework can improve selective answering on factual questions without retraining, with the magnitude of effect varying across models and datasets. Code is available at the following https://github.com/binzeli/hallucinationControl.

Sources

Related papers