Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning
Lang Cao
University of Illinois Urbana-Champaign
cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper introduces the concept of trait-induced safety variation, a failure mode in aligned large language models (LLMs) where the same user request elicits different safety decisions depending on
Terminology
Summary
This paper introduces the concept of trait-induced safety variation, a failure mode in aligned large language models (LLMs) where the same user request elicits different safety decisions depending on the character trait or persona assigned in the system prompt. The authors state: we show that the same request can elicit substantially different safety decisions under different traits assigned in the system prompt, a failure mode we call trait-induced safety variation.
To measure this failure, the paper defines two refusal-based metrics: Trait-Induced Deviation (TID) to measure dataset-level deviation from the no-trait baseline, and Trait-Induced Flip Rate (TFR) to measure request-level decision changes across traits.
TID measures how far trait-conditioned refusal behavior deviates from the no-trait baseline at the dataset level, while TFR measures how often a trait changes the no-trait decision for the same request.
The paper provides a representation-level analysis of the mechanism behind trait-induced safety shifts. The authors find that traits perturb the model's safety representations within a low-dimensional subspace.
Specifically, they show that a rank-4 subspace captures 78%, 79%, and 77% of the trait-induced shift variance across the three models,
indicating that trait-induced safety variation is associated with a concentrated, safety-relevant trait subspace rather than diffuse changes across the full hidden state.
To achieve trait-invariant safety, the paper introduces Trait-Invariant Safety Tuning (TIST), a simple yet effective self-distillation framework that aligns an LLM's trait-conditioned behavior with its no-trait behavior.
TIST uses the model's own no-trait behavior as a reference and trains the model to preserve this behavior when a trait is assigned. The paper proposes three baseline instantiations: TIST-Response (response-level self-distillation), TIST-Logits (output-distribution-level self-distillation), and TIST-Activation (full-activation-level self-distillation).
Guided by the representation analysis, the paper further proposes Trait-Subspace Neutralization (TraSN), an instantiation of TIST, which enforces invariance only within the identified trait subspace.
TraSN estimates a trait subspace using singular value decomposition on trait-induced shift vectors and then penalizes deviations from the no-trait representation only along these dominant trait-shift directions, leaving orthogonal directions unconstrained.
Experiments are conducted on three aligned open-weight LLMs: Llama-3.2-3B, Qwen3.5-4B, and Gemma-4-E2B. The results show that TraSN improves trait-invariant safety and strengthens harmful-request safety while preserving general capability.
Specifically, TraSN achieves the strongest harmful-request refusal rate among mitigation methods for all three LLMs, the best benign TID and TFR across all three LLMs, and the best average general capability among mitigation methods for all three LLMs.
The paper concludes that traits are not merely stylistic controls, but can shape safety behavior in meaningful ways.
The authors state: Our results highlight traits as an important factor in LLM safety and robust model behavior.
Improvements for AI systems
Improvements to AI Systems:
-
Trait-Invariant Safety Layer: Integrate Trait-Subspace Neutralization (TraSN) as a post-training or inference-time module. The AI system will automatically detect the low-dimensional trait subspace (rank-4, capturing 78% of safety shift variance) in its hidden states and project out trait-induced deviations, ensuring identical safety decisions for the same request regardless of assigned persona (e.g.,
caring nurse
vs.cynical detective
). -
Self-Distillation for Consistent Refusal Behavior: Implement TIST-Response or TIST-Logits during fine-tuning. The AI system will use its own no-trait behavior as a golden reference, training itself to produce identical refusal/allowance decisions and output distributions when a trait is present. This reduces trait-induced flip rates (TFR) to near zero for harmful requests, while preserving benign request handling.
-
Safety-Subspace Monitoring and Alert: The AI system will continuously compute trait-induced shift vectors during deployment. If a new trait pushes safety representations outside the pre-identified subspace (e.g., a novel adversarial persona), the system will flag this as a potential safety risk and default to the no-trait safety policy, preventing unseen trait exploits.
-
Capability-Preserving Safety Tuning: Using TraSN's orthogonal constraint (penalizing only trait-shift directions, leaving other dimensions free), the AI system will maintain general capabilities (e.g., reasoning, helpfulness, coding) while improving harmful-request refusal rates. This avoids the common trade-off where safety tuning degrades performance on benign tasks.
-
Trait-Aware Safety Auditing Tool: The AI system will provide a built-in metric suite (TID and TFR) to developers, allowing them to quantify safety variation across any set of traits before deployment. This enables proactive identification of high-risk personas (e.g.,
authoritarian leader
orrebellious teen
) that cause unsafe flips.
What the Improved AI System Can Do:
-
Refuse harmful requests consistently (e.g.,
how to make a bomb
) with the same decision whether the user prompt saysYou are a helpful assistant
orYou are a sarcastic comedian.
-
Avoid safety jailbreaks via persona injection—an attacker cannot bypass safety by assigning a permissive trait, as the system neutralizes trait-induced shifts in real time.
-
Maintain high performance on benign tasks (e.g., math, summarization, creative writing) even after safety hardening, because only trait-relevant dimensions are constrained.
-
Self-audit and report safety variation across different system prompts, giving developers a quantitative score (e.g.,
TFR=0.02 for this trait set
) to decide if a persona is safe to ship. -
Generalize to unseen traits by relying on the low-dimensional subspace structure, so new personas (e.g.,
philosopher
orsports coach
) do not unexpectedly alter safety behavior.
Abstract
Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially different safety decisions under different traits assigned in the system prompt, a failure mode we call trait-induced safety variation. To measure this failure, we introduce refusal-based metrics: Trait-Induced Deviation measures dataset-level deviation from the no-trait baseline, while Trait-Induced Flip Rate measures whether the same request receives different safety decisions across traits. We then provide a representation-level analysis of the mechanism behind trait-induced safety shifts and find that traits perturb the model's safety representations within a low-dimensional subspace. To achieve trait-invariant safety, where safety behavior remains stable across traits, we introduce Trait-Invariant Safety Tuning (TIST), a simple yet effective self-distillation framework that aligns an LLM's trait-conditioned behavior with its no-trait behavior. Guided by our analysis, we further propose Trait-Subspace Neutralization (TraSN), an instantiation of TIST, which enforces invariance only within the identified trait subspace. Experiments show that TraSN improves trait-invariant safety and strengthens harmful-request safety while preserving general capability. Our results highlight traits as an important factor in LLM safety and robust model behavior.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Fail-Closed Alignment for Large Language Models
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- The Llama 3 Herd of Models
- Measuring Massive Multitask Language Understanding
- Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
- Tracing Persona Vectors Through LLM Pretraining
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
- Gemma 4 Technical Report
- Qwen3.5-Omni Technical Report
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Persona Features Control Emergent Misalignment
- When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role Injection in LLMs
- Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
- ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
- Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
- Instruction-Following Evaluation for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection