When the Wrong Key Wins: Understanding and Detecting Hallucinations in LLMs
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-29
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models can hallucinate even when the knowledge required for a correct answer is already available.
Terminology
Abstract
Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a latent-key view of inference, where answer selection depends on competition among associations acquired during pretraining. We show that model predictions can be highly sensitive to individual query keywords, that these influential keywords exhibit entity-specific binding, and that their effects are systematically shaped by pretraining frequency. Multiple bindings can also compete and exhibit higher-order interactions within the same query. Based on this mechanism, we introduce a two-stage keyword-perturbation method for hallucination detection. By removing influential keywords and measuring how the model reorganizes its prediction, the method distinguishes errors caused by misleading key associations from correct decisions supported by diagnostic evidence. Across multiple models and benchmarks, perturbation provides a strong and transferable detection signal, reaching.910 AUROC on probe-known ScientistQA. Finally, we extend the same probabilistic framework to four hallucination regimes: knowledge deficit, wrong knowledge, context distraction, and unstable inference. Their operational distributions across benchmarks provide diagnostic context for why different detector families succeed in different settings.
Sources
- Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs
- Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers
- Gemini: A Family of Highly Capable Multimodal Models
- Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors
- Language Models (Mostly) Know What They Know
- WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences
- Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
- Sources of Hallucination by Large Language Models on Inference Tasks
- WebGPT: Browser-assisted question-answering with human feedback
- GPT-4 Technical Report
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
- Co-occurrence is not Factual Association in Language Models
- SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering