When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
cs.CR, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: Accepted to the Findings of EMNLP 2026
Code: https://github.com/Godblessmycode1/safety_routing
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
- GPT-4o System Card
- LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
- Proximal Policy Optimization Algorithms
- A StrongREJECT for Empty Jailbreaks
- The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
- Finetuned Language Models Are Zero-Shot Learners
- Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs