The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-28
Updated: 2026-10-03
Comments: 19 pages, 7 figures, 20 tables. Accepted to EMNLP 2026 (Main Conference). Code: https://github.com/js-lee-AI/LayerMix
Code: https://github.com/js-lee-AI/LayerMix
License: http://creativecommons.org/licenses/by/4.0/
The gist: Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures.
Terminology
Abstract
Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal overwhelmingly dominated by a single mean-shift component, and removing this direction collapses detection to chance. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full-dimensional classifiers, so apparent architectural complexity largely reflects high-dimensional covariance estimation difficulty rather than exploitable non-linearity. A simple L2-regularized logistic regression (0.952 AUROC) bounds or outperforms twelve controlled architectural alternatives, and our multi-layer aggregation exceeds CLAP cross-layer attention probing under matched paradigm. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle-layer performance without oracle access. Our claims characterize the geometry within the controlled paired-example paradigm. Our code is available at https://github.com/js-lee-AI/LayerMix.
Sources
- DRIFT: Detecting Representational Inconsistencies for Factual Truthfulness
- Discovering Latent Knowledge in Language Models Without Supervision
- INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
- H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
- The Llama 3 Herd of Models
- Linearity of Relation Decoding in Transformer Language Models
- HARP: Hallucination Detection via Reasoning Subspace Projection
- Mistral 7B
- Language Models (Mostly) Know What They Know
- Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- Neural Probe-Based Hallucination Detection for Large Language Models
- Teaching Models to Express Their Uncertainty in Words
- Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Steer LLM Latents for Hallucination Detection
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering