Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: 7 pages, 3 figures, 9 tables. Single-author preprint on white-box, single-pass hallucination detection using inter-layer activation divergence, cross-layer fusion, temporal drift, and calibrated risk scoring. Evaluated on TruthfulQA, HaluEval 2.0, and FaithDial with Llama-3, Qwen2.5, and Mistral backbones
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
- Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
- ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
- A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
- EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs
- The Geometry of Truth: Layer-wise Semantic Dynamics for Hallucination Detection in Large Language Models
- Automatic Layer Selection for Hallucination Detection
- Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
- Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations
- CrossHallu: Do Hallucination Signals Generalize Across Languages and Domains in Large Language Model's Internals?
- HALT: Hallucination Assessment via Log-probs as Time series
- Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering