The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Shared Category Geometry in Small Language Models
cs.LG
Submitted: 2026-07-18
Updated: 2026-09-02
Comments: Final version: extensively rewritten and restructured. Expanded with a replication campaign on a 14 model, 6 family and different architectures (including 2 MoEs), introducing a tiny library, for experiments. Code and data in 2 different repository at: https://github.com/Francesco-Marhel/
Code: https://github.com/Francesco-Marhel/TruthProbe
License: http://creativecommons.org/licenses/by/4.0/
The gist: Bürger et al.
Terminology
Abstract
Bürger et al. (2024) demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace. The truth value of a statement is linearly readable from a residual stream of language model, but it is not clear how much of that representation fits on a single direction, which component builds it, or what it is made of. We conducted a study based on these questions, with one instrument: a training-free axis, the dominant direction of the singular value decomposition (SVD) of hidden-state differences over true/false minimal pairs, identified without labels up to one global sign. Extensive evaluation across 14 models from 6 diverse architectural families (including MoE), read and extract at cost O(d) per token. We close with a pre-registered prediction on whether the arrangement extends to categories whose truth is computed rather than retrieved.
Sources
- Discovering Latent Knowledge in Language Models Without Supervision
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Representation Engineering: A Top-Down Approach to AI Transparency
- Truth is Universal: Robust Detection of Lies in LLMs
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Toy Models of Superposition
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
- LLM Hallucination Detection: A Fast Fourier Transform Method Based on Hidden Layer Temporal Signals
- The Platonic Representation Hypothesis
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks