Linear Separability of Activation Representations after Supervised Fine-Tuning on Incorrect Responses: A Study of Synthetic Dishonesty in Large Language Models

arXiv:2605.30381 · cs.LG, cs.AI · Submitted 2026-05-28 · Read on arXiv

cs.LG, cs.AI

Submitted: 2026-05-28

Updated: 2026-09-05

Code: https://github.com/vzm1399/llm-dishonesty-representations

License: http://creativecommons.org/licenses/by/4.0/

The gist: When a language model is fine-tuned to produce systematically incorrect responses, does this training leave a structured, linearly recoverable trace in its internal activations? We study this

Terminology

Abstract

When a language model is fine-tuned to produce systematically incorrect responses, does this training leave a structured, linearly recoverable trace in its internal activations? We study this question in a controlled model-organism setting using five transformer architectures spanning 1.4 to 9 billion parameters. For each model, an "honest" and a "dishonest" LoRA fine-tuned variant are constructed using identical question distributions but correct versus plausible-but-incorrect answers. Linear probes reach near-ceiling separability (AUC >= 0.9997) within the first few layers in four of five architectures. This separability transfers from TruthfulQA to held-out MMLU subjects for four models, but substantially less so for Pythia-1.4B. Six geometric analyses reveal an architectural dichotomy in representation structure. A control experiment further shows that LoRA fine-tuning itself leaves a strong activation fingerprint that is largely dissociable from the honesty-related signal, independent of fine-tuning data content. Two supplementary experiments qualify these findings: the identified direction transfers poorly to paraphrased questions, with AUC approaching chance, and norm-calibrated activation steering provides no reliable evidence of a causal role in generation, partly because stronger interventions disrupt output fluency. These results characterize what supervised fine-tuning on incorrect responses induces in activation space while clarifying important limitations on interpretation and causal significance.

Related papers