Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 18 pages, 6 figures. Accepted to the Main Conference of EMNLP 2026
Code: https://github.com/hjhj97/Empathy-Is-Steerable-but-Multi-Axial
License: http://creativecommons.org/licenses/by/4.0/
The gist: Activation steering has been used to control traits such as honesty, refusal, and sycophancy, yet supportive empathy is evaluated along multiple dimensions that need not correspond to independently
Terminology
Abstract
Activation steering has been used to control traits such as honesty, refusal, and sycophancy, yet supportive empathy is evaluated along multiple dimensions that need not correspond to independently controllable activation directions. Using the EPITOME framework, which decomposes supportive empathy into Emotional Reactions, Interpretations, and Explorations, we study three instruction-tuned LLMs and ask whether candidate directions derived from these labels produce distinguishable intervention effects or instead share structure, and how persona prompts interact with those directions. We find that contrastive activation addition yields a stable middle-layer intervention that consistently shifts the EPITOME proxy scores across models, moving empathy analysis beyond response-level scoring. However, the recovered directions are only partially separable: steering one direction induces off-target shifts, and hand-crafted prompting shifts the empathy profile rather than isolating a single dimension. Persona prompts substantially change EPITOME scores, but a paired activation-shift decomposition shows that the recovered subspace captures only approximately 3 percent of persona-induced squared activation-shift magnitude at layer 15. Under this EPITOME-based definition, expressed empathy is steerable but multi-axial, and controlling persona-conditioned empathy requires targeting structure beyond individual mechanism directions.
Sources
- Refusal in Language Models Is Mediated by a Single Direction
- The Llama 3 Herd of Models
- PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
- HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
- Mistral 7B
- Large Language Models Produce Responses Perceived to be Empathic
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
- Knowledge Vector of Logical Reasoning in Large Language Models
- Qwen2 Technical Report
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering