Where is the Mind? Persona Vectors and LLM Individuation
cs.CL, cs.AI
Submitted: 2026-04-18
Updated: 2026-09-09
Code: https://github.com/bepierre/whereis-the-mind-mini-experiments
License: http://creativecommons.org/licenses/by/4.0/
The gist: The individuation problem for large language models asks which entities associated with them, if any, should be identified as minds.
Terminology
Abstract
The individuation problem for large language models asks which entities associated with them, if any, should be identified as minds. We approach this problem through mechanistic interpretability, engaging in particular with recent empirical work on persona vectors, persona space, and emergent misalignment. We argue that three views are the strongest candidates: the virtual instance view and two new views we introduce, the (virtual) instance-persona view and the model-persona view. First, we argue for the virtual instance view on the grounds that attention streams sustain quasi-psychological connections across token-time. Then we present the persona literature, organised around three hypotheses about the internal structure underlying personas in LLMs, and show that the two persona-based views are promising alternatives.
Sources
- Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
- Constitutional AI: Harmlessness from AI Feedback
- Going Whole Hog: A Philosophical Defense of AI Cognition
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- The Artificial Self: Characterising the landscape of AI identity
- One-shot Optimized Steering Vectors Mediate Safety-relevant Behaviors in LLMs
- Who's asking? User personas and the mechanics of latent misalignment
- Predictive Minds: LLMs As Atypical Active Inference Agents
- Linear representations in language models can change dramatically over a conversation
- Taking AI Welfare Seriously
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Palatable Conceptions of Disembodied Being
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
- Convergent Linear Representations of Emergent Misalignment
- Model Organisms for Emergent Misalignment
- Persona Features Control Emergent Misalignment
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering