"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-09
Updated: 2026-08-09
Comments: Accepted to COLM 2026 Workshop on Efficient Reasoning and KONVENS 2026 First Workshop on Evaluating LLMs for Specialized Domains (Eval4SD)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves.
Terminology
Abstract
Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or self-knowledge of the models, yet what drives them is not well understood. Are the models telling us about themselves or rather how they are deployed? In this work, we show that the chat template works like a switch - when present, it turns this disclaimer voice up and experiential voice like "I feel" down, across 8 popular open-source instruct models up to 9B parameters in size. And conversely when the chat template is not present, it turns the disclaimer voice down and experiential voice up. Inside the activations of 3 models, we find a direction that steers this behavior. Removing the direction in the model's activation space turns disclaimer voice down and adding it turns it up, while a random direction of the same size has little effect. We find that instruct models without chat template, when we add the disclaimer direction to them, disclaim like the template was there. Since the chat template controls the disclaimer voice of LLMs, then researchers studying self-reports or introspection of models might have a confound they need to control for. Our results show that there is a direction they can use to steer this voice. More broadly, our work shows that what models say about themselves is not a fact about them. What they say doesn't come only from weights, but it is partially set by the chat template, and because of that a model's self-description shouldn't be treated literally.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Large Language Models Report Subjective Experience Under Self-Referential Processing
- Taken out of context: On measuring situational awareness in LLMs
- Tell me about yourself: LLMs are aware of their learned behaviors
- Looking Inward: Language Models Can Learn About Themselves by Introspection
- Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
- Consciousness in Artificial Intelligence: Insights from the Science of Consciousness
- Self-Cognition in Large Language Models: An Exploratory Study
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Does It Make Sense to Speak of Introspection in Large Language Models?
- A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems
- The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- There Is More to Refusal in Large Language Models than a Single Direction
- Language Models (Mostly) Know What They Know
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- Emergent Introspective Awareness in Large Language Models
- Taking AI Welfare Seriously
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Towards Evaluating AI Systems for Moral Status Using Self-Reports
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks