The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Shared Category Geometry in Small Language Models

arXiv:2607.16741 · cs.LG · Submitted 2026-07-18 · Read on arXiv

cs.LG

Submitted: 2026-07-18

Updated: 2026-09-02

Comments: Final version: extensively rewritten and restructured. Expanded with a replication campaign on a 14 model, 6 family and different architectures (including 2 MoEs), introducing a tiny library, for experiments. Code and data in 2 different repository at: https://github.com/Francesco-Marhel/

Code: https://github.com/Francesco-Marhel/TruthProbe

License: http://creativecommons.org/licenses/by/4.0/

The gist: Bürger et al.

Terminology

Abstract

Bürger et al. (2024) demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace. The truth value of a statement is linearly readable from a residual stream of language model, but it is not clear how much of that representation fits on a single direction, which component builds it, or what it is made of. We conducted a study based on these questions, with one instrument: a training-free axis, the dominant direction of the singular value decomposition (SVD) of hidden-state differences over true/false minimal pairs, identified without labels up to one global sign. Extensive evaluation across 14 models from 6 diverse architectural families (including MoE), read and extract at cost O(d) per token. We close with a pre-registered prediction on whether the arrangement extends to categories whose truth is computed rather than retrieved.

Sources

Related papers