Decoupled Alignment for Robust Plug-and-Play Adaptation
cs.CL, cs.AI, cs.CR
Submitted: 2024-06-03
Updated: 2026-08-26
Comments: Conference on Language Modeling (COLM) 2026
Code: https://github.com/NWULIST/DAPA
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback.
Terminology
Abstract
We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust plug-and-play approach to prevent shadow alignment when models are adapted to downstream tasks. Specifically, we leverage knowledge distillation to extract alignment signals from well-aligned LLMs and inject them into shadow-aligned models via model fusion, enabling plug-and-play alignment correction. In our methodology, we employ delta debugging to identify the critical components of knowledge necessary for effective distillation. On the harmful question dataset, our method significantly enhances the average defense success rate by approximately 14.42%, reaching as high as 51.39% across 17 influenced LLMs, without compromising performance. Our code is available at https://github.com/NWULIST/DAPA.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen Technical Report
- Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- The Llama 3 Herd of Models
- Gaussian Error Linear Units (GELUs)
- Nonparametric Modern Hopfield Models
- GPT-4o System Card
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- GPT-4 Technical Report
- Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action
- Instruction Tuning with GPT-4
- GLU Variants Improve Transformer
- Gemma: Open Models Based on Gemini Research and Technology
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen2.5 Technical Report
- Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
- DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering