The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
cs.LG, cs.AI
Submitted: 2026-09-07
Updated: 2026-09-25
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi et al., 2024), and activation-space probes (Arditi
Terminology
Abstract
Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi et al., 2024), and activation-space probes (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. We give this fragility one geometric explanation and trace it to when, during pretraining, safety can take hold. We measure the safety update Δ= W safe - W base against the curvature of the model's capabilities (the empirical Fisher of a capability loss). Post-hoc safety consistently lands in a suppression regime: Δ is nearly orthogonal to the capability directions, and its small in-subspace part concentrates on a few high-curvature ones. The update is thin but sharp, a refusal gate laid over intact capabilities rather than erasure of them. A kernel-immobility lemma explains why such an update can only mask a capability, not remove it, so a little benign fine-tuning restores it: 100 steps of benign fine-tuning collapse refusal on Qwen-2.5-7B and Llama-3-8B Instruct at preserved capability, a signature that replicates across five model families. Following the account into pretraining, a 267-checkpoint sweep of OLMo-2-1B (OLMo et al., 2025) shows the substrate that safety engages emerging in a sharp transition between roughly 6B and 60B pretraining tokens. We then use the account constructively: models trained from scratch with safety co-training spread continuously across pretraining reach 87 to 98% refusal whose post-attack level holds at 84 to 91% at every scale, an erosion of 2 to 14 pp against 35 to 38 pp for post-hoc installs, at capability matched or better than an LM-only baseline and holding from 410M to 6.9B, whereas a compute-matched windowed schedule installs no lasting refusal. Persistence of the safety signal across pretraining, not its timing, is what buys attack robustness.
Sources
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Training Verifiers to Solve Math Word Problems
- Investigating Critical Period Effects in Language Acquisition through Neural Language Models
- The Llama 3 Herd of Models
- Alignment faking in large language models
- Studying Large Language Model Generalization with Influence Functions
- Gradient Descent Happens in a Tiny Subspace
- Measuring Massive Multitask Language Understanding
- Training Compute-Optimal Large Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
- Eight Methods to Evaluate Robust Unlearning in LLMs
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks