The Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 12 pages, 4 tables. Accepted to AACL-IJCNLP 2026 (main conference)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Three recent results describe what look like unrelated LLM reliability problems.
Terminology
Abstract
Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show reasoning RL collapses tool-reliability representations. Suleymanov et al. (2026) show that under safety-constrained generation, large models rewrite flagged spans while small models truncate. Bastounis et al. (2024) prove any consistent-reasoning system without an implicit "I don't know" function must hallucinate infinitely often on broad problem classes. We argue these findings converge on a single intervention: calibrated abstention is what each independently identifies as the missing capability, even though the unavailability they document, a capability gap, a policy gap, and a recursion-theoretic gap, has a different source in each case. Honesty post-training has narrowed the gap in deployed models, but principled closure of the class Bastounis identifies requires a calibrated abstention function whose training signal at the leaderboard level is absent: dominant benchmarks assign zero reward to decline, so the leaderboard gradient that would select for the function does not exist. We propose four changes to evaluation: triple-scoring, abstention-rate reporting, capability-stratified evaluation, and mandatory calibration metrics. Benchmark reform is necessary, not sufficient, for closing the gap the theorem identifies.
Sources
- On the consistent reasoning paradox of intelligence and optimal trust in AI: The power of 'I don't know'
- Evaluating Large Language Models Trained on Code
- The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
- Training Verifiers to Solve Math Word Problems
- Measuring Massive Multitask Language Understanding
- AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
- Language Models (Mostly) Know What They Know
- Why Language Models Hallucinate
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Holistic Evaluation of Language Models
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Beyond Refusal: Probing the Limits of Agentic Self-Correction for Semantic Sensitive Information
- Measuring short-form factuality in large language models
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- Know Your Limits: A Survey of Abstention in Large Language Models
- The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks