When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
cs.CL, cs.AI, cs.LG
Submitted: 2025-11-10
Updated: 2026-09-11
Code: https://github.com/KellerJordan/modded-nanogpt
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite substantial advances, large language models (LLMs) continue to exhibit hallucinations, generating plausible yet incorrect responses.
Terminology
Abstract
Despite substantial advances, large language models (LLMs) continue to exhibit hallucinations, generating plausible yet incorrect responses. In this paper, we highlight a critical yet previously underexplored class of hallucinations driven by spurious correlations -- superficial but statistically prominent associations between features (e.g., surnames) and attributes (e.g., nationality) present in the training data. We demonstrate that these spurious correlations induce hallucinations that are confidently generated, immune to model scaling, evade current detection methods, and persist even after refusal fine-tuning. Through systematically controlled synthetic experiments and empirical evaluations on state-of-the-art open-source and proprietary LLMs (including GPT-5), we show that existing hallucination detection methods, such as confidence-based filtering and inner-state probing, fundamentally fail in the presence of spurious correlations. Our theoretical analysis further elucidates why these statistical biases intrinsically undermine confidence-based detection techniques. Our findings thus emphasize the urgent need for new approaches explicitly designed to address hallucinations caused by spurious correlations.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Reading Wikipedia to Answer Open-Domain Questions
- Can AI Assistants Know What They Don't Know?
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
- Deep Think with Confidence
- Beyond ReLU: How Activations Affect Neural Kernels and Random Wide Networks
- SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
- Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
- ConfRAG: Confidence-Guided Retrieval-Augmenting Generation
- Why Language Models Hallucinate
- Adam: A Method for Stochastic Optimization
- HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
- How Pre-trained Language Models Capture Factual Knowledge? A Causal-Inspired Analysis
- Out-of-Distribution Detection Methods Answer the Wrong Questions
- DeepSeek-V3 Technical Report
- Uncertainty Estimation in Autoregressive Structured Prediction
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering