Testing the Limits of Truth Directions in LLMs
cs.CL, cs.AI
Submitted: 2026-04-04
Updated: 2026-09-11
Comments: BlackboxNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction.
Terminology
Abstract
Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction. Previous studies have argued that these directions are universal in certain aspects, while more recent work has questioned this conclusion drawing on limited generalization across some settings. In this work, we identify a number of limits of truth-direction universality that have not been previously understood. We first show that truth directions are highly layer-dependent, and that a full understanding of universality requires probing at many layers in the model. We then show that truth directions depend heavily on task type, emerging in earlier layers for factual and later layers for reasoning tasks; they also vary in performance across levels of task complexity. Finally, we show that model instructions can affect truth directions; simple correctness evaluation instructions significantly affect the geometry and the generalization ability of the truth probes. Our findings indicate that universality claims for truth directions are more limited than previously known, with significant differences observable for various model layers, task difficulties, task types, and prompt templates.
Sources
- Understanding intermediate layers using linear classifier probes
- The Internal State of an LLM Knows When It's Lying
- The Geometries of Truth Are Orthogonal Across Tasks
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Discovering Latent Knowledge in Language Models Without Supervision
- Truth is Universal: Robust Detection of Lies in LLMs
- No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
- The Llama 3 Herd of Models
- LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- Adam: A Method for Stochastic Optimization
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Gemma: Open Models Based on Gemini Research and Technology
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering