Measuring the Assistant's Harmlessness Preferences on the User Turn
cs.CL
Submitted: 2026-09-20
Updated: 2026-09-20
License: http://creativecommons.org/licenses/by/4.0/
The gist: Post-training turns a general next-token predictor into a chat model with a persistent assistant persona.
Terminology
Abstract
Post-training turns a general next-token predictor into a chat model with a persistent assistant persona. If that persona is a character the model plays only on its own turns, its preferences should govern what the assistant says, not what the model predicts other speakers will say. We test this boundary and find that it does not hold: a safety-relevant preference of the assistant---for harmless over harmful tasks---shapes the model's predictions even on the user's turn, where the assistant is not the one speaking. We find that this preference is small or near-zero in pretrained base models, that it emerges through post-training, replicated across open-weight model families, grows with scale, and can be moved by narrow finetuning that never touches user turns. We claim that this is evidence that post-training does not merely install a shallow assistant persona, but instead generalises beyond just the local assistant turn, into the model's representation of the user.
Sources
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
- Mechanisms of Introspective Awareness
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- From Simulation to Enaction: Post-trained language models recognize and react to their own generations
- Gemma 3 Technical Report
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
- The Llama 3 Herd of Models
- Qwen2.5 Technical Report
- Convergent Linear Representations of Emergent Misalignment
- Olmo 3
- Persona Features Control Emergent Misalignment
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering