HalluTracer: Pre-Decoding Truthfulness Prediction via Depth-Averaged Probe-Logit
cs.CL, cs.AI
Submitted: 2026-08-17
Updated: 2026-09-19
License: http://creativecommons.org/licenses/by/4.0/
The gist: Internal-state probes enable truthfulness prediction before a large language model generates an answer.
Terminology
Abstract
Internal-state probes enable truthfulness prediction before a large language model generates an answer. When detectors change both the layers they read and the rules used to combine them, the source of improved prediction becomes difficult to identify. We separate these choices and find that retaining more layers improves prediction even under fixed equal weighting. An exact Fisher-ratio decomposition explains why the additional benefit of linear reweighting is limited on these probe scores: information inlayer-wise differences largely overlaps with that captured by the depth mean. Estimating additional weights can then offset this small benefit when the data used to fit them are limited. These findings motivate our proposed method HalluTracer, which averages layer-wise probe logits to predict truthfulness before decoding. Across six models and four benchmarks, including TruthfulQA, HalluTracer achieves the highest area under the receiver operating characteristic curve (AUROC) in 23 of 24 model--benchmark pairs among the compared methods. The results support using evidence from across the network without requiring a correspondingly more flexible aggregation rule, clarifying the distinct roles of layer selection and weighting in pre-decoding detection.
Sources
- The Internal State of an LLM Knows When It's Lying
- Constitutional AI: Harmlessness from AI Feedback
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Are Your Agents Upward Deceivers?
- Language Models (Mostly) Know What They Know
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- The Geometry of Truth: Layer-wise Semantic Dynamics for Hallucination Detection in Large Language Models
- Training language models to follow instructions with human feedback
- Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering