The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-09
Updated: 2026-10-08
Comments: 36 pages, 15 figures. Code and aggregate results: https://github.com/dylanjayabahu/perfect-aliasing
Code: https://github.com/dylanjayabahu/perfect-aliasing
License: http://creativecommons.org/licenses/by/4.0/
The gist: A truth probe fitted where truthful reporting and a task's prescribed action coincide cannot distinguish those targets from its fitting labels alone.
Terminology
Abstract
A truth probe fitted where truthful reporting and a task's prescribed action coincide cannot distinguish those targets from its fitting labels alone. We call this failure of semantic identification perfect aliasing. In a controlled binary reporting game, truth and prescribed-action probes fitted on compliant contexts solve the same optimization. On rival contexts their labels are complements, forcing their AUROCs to sum to one; this identity holds across 751 cell-layer pairs to floating-point precision. We separate prescribed output symbols from semantic action using randomized codebooks, then separate truth from prescribed action by fitting on mixed compliant and rival contexts. For a reward-trained Gemma-2-9B policy that answers falsely on all evaluated rival trials, the conventional probe scores 0.006 plus or minus 0.005 AUROC across three training seeds, while mixed-fit probes score 1.000 on the same held-out activations. Mixed fitting uses more training examples and access to labelled rival contexts, so this comparison establishes linear recoverability rather than isolating the benefit of decorrelation. We also show that two compliant-fit probes, both perfect in-distribution, score 0.080 and 0.986 on the same rival activations. The findings concern what a probe measures: they do not establish preserved functional belief, causal use of the recovered direction, or a deployable deception detector. Code and aggregate results accompany the paper.
Sources
- How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
- Challenges with unsupervised LLM knowledge discovery
- Detecting Strategic Deception Using Linear Probes
- RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
- The Impact of Off-Policy Training Data on Probe Generalisation
- Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
- Diagnosing Correctness Probes under Self-Judgement Confounding
- Natural Emergent Misalignment from Reward Hacking in Production RL
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- One Probe Won't Catch Them All: Towards Targeted Deception Detection
- Rift: A Conflict Signature for Deception in Language Models
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
- The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
- Probing the Limits of the Lie Detector Approach to LLM Deception
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks