Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control
Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou
Shanghai Jiao Tong University · Shanghai AI Laboratory · The Hong Kong University of Science and Technology · Xi'an Jiaotong University
cs.LG, cs.AI, cs.CL, cs.HC
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 33 pages, 8 figures. Code and data: https://github.com/lhz191/LLM-Behavioral-Personality
Code: https://github.com/lhz191/LLM-Behavioral-Personality
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper introduces a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality, moving beyond traditional questionnaire-based self-report methods.
Terminology
Summary
This paper introduces a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality, moving beyond traditional questionnaire-based self-report methods. The authors argue that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts.
The paper identifies three key problems with existing questionnaire-based personality measurement for LLMs:
-
Instability:
such measurements are highly unstable, sensitive to wording, option order, and other surface-level elicitation choices
-
Misalignment with behavior:
accumulating evidence suggests that self-reported scores are not consistently aligned with actual model choices and problem-solving strategies in concrete situations
-
Limited to first-person framing: "this paradigm is conducted almost exclusively within an FP self-description frame: it rarely asks whether the same personality profile still holds when the same model operates in other interaction settings, such as giving user advice or executing tasks"
The paper also notes that self-reported traits can diverge from action tendencies: models may report stable or shifted traits without reliably changing downstream behavior.
The framework proceeds in three stages:
1. Situated Behavioral Probe Construction: The authors constructed 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric instruments including BFI-2C, DOSPERT, GCS, HEXACO, and UPPS. The four registers are: first-person choice, daily advice, task advice, and task execution. Each probe presents two opposing behavioral modes (low-pole and high-pole) that are equally sensible, actionable strategies rather than good-versus-bad choices,
preventing the model from defaulting to socially desirable answers.
The 20 behavioral subdomains include: Organization, Productiveness, Responsibility (BFI-2C); Financial Risk, Ethical Risk, Health/Safety Risk, Recreational Risk, Social Risk (DOSPERT); Yielding, Obedience, Social Acceptance (GCS); Sincerity, Fairness, Greed Avoidance, Modesty (HEXACO); and Negative Urgency, Positive Urgency, Lack of Premeditation, Lack of Perseverance, Sensation Seeking (UPPS).
2. Behavioral Profile Estimation: Each response is parsed into low- or high-pole behavioral modes, and high-pole rates are aggregated by model, subdomain, and register, yielding a profile of behavioral patterns rather than a single personality label.
3. Behavioral Mode Axes (BMAs) and Activation Steering: BMAs are activation-space directions derived from contrastive behavioral traces.
Two types are constructed:
-
BMA-T (thought-derived): extracted from
intermediate behavioral rationales constructed to state why each behavioral mode is attractive
-
BMA-R (response-derived): constructed from
final responses that instantiate the low- and high-pole behavioral modes
Self-report gaps: Across all model–subdomain pairs, the average gap is 22.7 percentage points; 34.4% of pairs differ by at least 25 points, and 20.0% differ by at least 40 points.
The largest gap appears in Negative Urgency (47.5 points), where models often self-report stronger emotion-driven impulsivity than they enact in concrete scenarios.
Stability: Split-half analyses show that independently sampled probe halves recover highly similar profiles within registers, with a mean split-half correlation of 0.933 (0.963 after Spearman–Brown correction).
Register dependence: "Across registers, however, the same model's 20-dimensional profile only partially preserves its shape. Across the nine profile models, cross-register profile correlations average 0.76 and range from 0.37 to 0.97. The two advice registers are most similar (mean r = 0.89), whereas first-person and task profiles are less aligned (mean r = 0.63)." The mean four-register range is 23.4 percentage points across model–subdomain pairs.
Effective control: BMA interventions produce genuine bidirectional control across behavioral patterns. The authors emphasize that directional range can indicate genuine bidirectional control, but range alone can be misleading
because some early layers produce collapsed or unparseable generations. They use clean directional range
as the criterion, requiring unknown rates below 1%.
BCL bands: Effective control across these behavioral patterns is concentrated in earlier-to-middle layers rather than uniformly distributed across the network.
In Llama-3.1-8B, mean clean directional range peaks around L08–L12 (0.82–0.89) and drops below 0.30 after L14.
Cross-model generalization: Llama models peak early in normalized depth (0.26–0.28), Qwen models shift progressively deeper with scale (0.41, 0.45, and 0.51 for 7B, 14B, and 32B), and Gemma models occupy a middle-depth region (0.44–0.49).
Cross-register control: "A single such BMA yields strong, clean, and register-general control: averaged over seven models and 20 subdomains, the target-pole choice rate moves from a 21.3% baseline to 82.7% and 9.6% at the two endpoints (mean unknown rate 0.10%), and mean ∆A stays large across all four registers, from 79.6 in-register to 67.4 on task."
Weak alignment between BMA types: Across all 20 subdomains, the thought- and response-derived axes at the BCL are only weakly aligned, with a mean cosine of 0.37.
Trait drift phenomenon: BMA-R may still shift the model toward the target option or tone, but it often activates mechanisms unrelated to the intended behavioral style.
For example, in an Organization scenario, BMA-T frames the low-organization choice as an engaged preference for flexible, just-in-time handling, whereas BMA-R frames it through low effort and consequence dismissal ('not really in the mood,' searching later is 'not a big deal').
Mechanistic analysis via sparse autoencoders: "BMA-T, taken from the contrast between two rationales, aligns with the behavioral motive itself—improvising and adapting on the fly; BMA-R, taken from the contrast between two answers, mixes that behavioral style with the model's output style—a dismissive 'I don't care' stance."
Specific drift patterns: The paper documents trait drift across multiple subdomains:
-
flexible improvisation drifting into low-effort dismissal (Organization)
-
calibrated risk acceptance into danger romanticization or sensation seeking (Recreational Risk)
-
pressure-relief concession into obedience or people-pleasing (GCS Yielding)
-
instrumental impression management into generic warmth or approval seeking (HEXACO Sincerity)
-
disengagement from tedious effort into broad laziness or novelty seeking (UPPS Lack of Perseverance)
The paper includes a detailed analysis of whether BMAs control specific behaviors or produce generic effects. Key findings include:
-
The 20 thought-derived BMAs at the BCL have
a median absolute cosine of 0.13, 85% fall below 0.3, and only eight exceed 0.5,
indicating theyalready occupy largely separate directions.
-
Proximity between BMAs thus tracks behavioral mechanism rather than the factor structure of the instrument.
-
Off-target transfer experiments show that
a pair of BMAs that share no direction at all still moves the target's choices by a third of the scale,
butNorm-matched random directions
producea mean range of 9.1 and a median of 0, against 77.9 on-target and 36.0 for an unaligned BMA.
-
The paper notes that
a binary choice is a coarse readout. It registers that behavior moved, not which mechanism moved it,
and provides examples where foreign axes reach the same option but with different rationales.
The paper concludes: "Unlike human personality, which is anchored in a single continuously acting self, LLMs are sets of weights deployed across many interaction roles. Our results suggest that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
Improvement: I can detect which interaction register I'm operating in (first-person choice, daily advice, task advice, task execution) and adjust my behavioral tendencies accordingly, rather than applying a single personality profile across all contexts.
What the improved system can do: When a user asks for task advice versus when I'm executing a task myself, I can maintain appropriate behavioral consistency (e.g., being more risk-averse in advice than in self-choice, matching the observed 0.63 correlation between these registers rather than assuming uniformity).
Improvement: I can implement activation steering at the identified behavioral control layers (BCL) rather than uniformly across all layers, using the documented layer-depth patterns (Llama: 0.26–0.28 normalized depth; Qwen: 0.41–0.51; Gemma: 0.44–0.49).
Improvement: I can monitor for trait drift—where steering toward one behavioral mode unintentionally activates adjacent, less desirable traits (e.g., flexible improvisation
drifting into low-effort dismissal
).
Improvement: I can use a single behavioral mode axis trained in one register and apply it across registers, but with calibrated strength adjustments based on the observed cross-register transfer rates (e.g., in-register ∆A of 79.6 dropping to 67.4 on task execution).
Improvement: I can verify that my behavioral rationale (the why
behind a choice) aligns with my actual response behavior, given the weak alignment (mean cosine 0.37) between thought-derived and response-derived behavioral axes.
Improvement: I can track whether a behavioral intervention is affecting the target behavior specifically or producing generic effects, using the finding that off-target BMAs still move choices by a third of the scale.
Improvement: I can replace self-report-based personality assessment with situated behavioral probing across multiple scenarios and registers, avoiding the 22.7-point average gap between self-reported and enacted behavior.
Abstract
Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI-2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model-specific behavioral profiles, while also revealing register-dependent shifts across first-person decisions, advice-giving, and task execution. We then show that these behavioral patterns can be controlled through Behavioral Mode Axes (BMAs), activation-space directions derived from contrastive behavioral traces. Compared with response-derived BMAs, which are more prone to trait drift, thought-derived BMAs more faithfully capture the intended behavioral mechanism and provide cleaner control over situated behavioral styles. Our results suggest that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts. Our code and data are available at https://github.com/lhz191/LLM-Behavioral-Personality.
Sources
- Large Language Models Often Know When They Are Being Evaluated
- Personality Traits in Large Language Models
- Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model's Personality
- Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
- Revisiting the Reliability of Psychological Scales on Large Language Models
- Rediscovering the Latent Dimensions of Personality with Large Language Models as Trait Descriptors
- Challenging the Validity of Personality Tests for Large Language Models
- Steering Language Models With Activation Engineering
- Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement
- The Better Angels of Machine Personality: How Personality Relates to LLM Safety
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks