Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks
cs.LG, cs.AI, cs.CR
Submitted: 2026-08-26
Updated: 2026-09-03
License: http://creativecommons.org/licenses/by/4.0/
The gist: Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse.
Terminology
Abstract
Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace shared across harmful prompts: ablating those directions suppresses refusals while largely preserves other model capabilities. Yet it remains unclear why safety-critical features in a wide range of models emerge in a concentrated, low-dimensional structure. In a case study of OLMo-2-0425-1B-Instruct we find that the refusal geometry reflects refusal training: activation updates resulting from refusal-completion first-token losses explain the resulting refusal direction and refusal subspace. We study refusal directions through the training dynamics across refusal datasets and reveal that their brittleness is associated with repetitive refusal starts, which in turn is linked to concentration of gradients and refusal features in a low-dimensional subspace. Across frozen-model analyses and controlled synthetic fine-tuning, we find evidence of a hardening lever: diverse refusal starts can raise stable ranks of gradients and activation changes, making refusals harder to remove with a vector ablation attack.
Sources
- Plug and Play Language Models: A Simple Approach to Controlled Text Generation
- Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
- Constitutional AI: Harmlessness from AI Feedback
- Training language models to follow instructions with human feedback
- Extracting Latent Steering Vectors from Pretrained Language Models
- Frontier AI Regulation: Managing Emerging Risks to Public Safety
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Will releasing the weights of future large language models grant widespread access to pandemic agents?
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Jailbroken: How Does LLM Safety Training Fail?
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Refusal in Language Models Is Mediated by a Single Direction
- Stealing Part of a Production Language Model
- Steering Llama 2 via Contrastive Activation Addition
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep
- Function Vectors in Large Language Models
- Steering Language Models With Activation Engineering
- Steering Large Language Model Activations in Sparse Spaces
- Toward universal steering and monitoring of AI models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks