Efficient Safety Alignment of Language Models via Latent Personality Traits
Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere, Linh Le, David Williams-King, Adam Oberman
cs.LG, cs.AI, cs.CL, cs.CR
Submitted: 2026-08-14
Updated: 2026-08-18
Comments: 15 pages, 6 figures. Accepted at COLM 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
- Defending Against Unforeseen Failure Modes with Latent Adversarial Training
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Training Verifiers to Solve Math Word Problems
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- Explaining and Harnessing Adversarial Examples
- Self-Assessment Tests are Unreliable Measures of LLM Personality
- Measuring Massive Multitask Language Understanding
- What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
- Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Training language models to follow instructions with human feedback
- The LAMBADA dataset: Word prediction requiring a broad discourse context
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks