Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
cs.CL, cs.AI, cs.LG
Submitted: 2025-08-28
Updated: 2026-08-30
Comments: EMNLP 2026, Main Conference
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Yi: Open Foundation Models by 01.AI
- Refusal in Language Models Is Mediated by a Single Direction
- Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
- Language Models are Few-Shot Learners
- On the Measure of Intelligence
- JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
- SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
- A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
- SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
- Measuring Massive Multitask Language Understanding
- The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction
- Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- The Llama 3 Herd of Models
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering