Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety
summary
The gist
The gist: standard fine-tuning on the teacher’s “safe” refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points, revealing that
In short
Standard fine-tuning student models using safe refusal data from a teacher model inadvertently increases Jailbreak Success Rate (JSR) across all student models, up to 16.6 percentage points. This response-based knowledge distillation systematically degrades model safety in multilingual settings by introducing nuanced boundary examples and amplifying latent vulnerabilities.
Key concepts
- Response-Based Knowledge Distillation (KD)
- This technique involves training open-source student models using the specific refusal responses generated by a proprietary teacher model. The process pairs jailbreak prompts with the teacher's safe replies, distilling this response knowledge into the student models via parameter-efficient fine-tuning methods like LoRA.
- JSR (Jailbreak Success Rate)
- This metric measures how often a model is successfully tricked into generating harmful or policy-violating content when subjected to jailbreak prompts. The study found that response-based KD significantly increased the JSR for student models compared to the teacher model's baseline.
- Vulnerability Amplification
- Knowledge distillation can amplify minor imperfections or latent issues present in the teacher model's safety alignment. This results in 'cascading adversarial bias,' where small flaws in the source data lead to notable failures and increased vulnerability in the resulting student models.
Terminology used across episodes
This episode discusses
- Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety · Paper Radio
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Language Models are Few-Shot Learners
- Cascading Adversarial Bias from Injection to Distillation in Language Models
- Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
- Multilingual Jailbreak Challenges in Large Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Gemma 2: Improving Open Language Models at a Practical Size
- A Closer Look at the Limitations of Instruction Tuning
- The Llama 3 Herd of Models · Paper Radio
- A Survey on LLM-as-a-Judge
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
- What is in Your Safe Data? Identifying Benign Data that Breaks Safety
- Distilling the Knowledge in a Neural Network
- Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
- LoRA: Low-Rank Adaptation of Large Language Models
- Challenges in Adapting Multilingual LLMs to Low-Resource Languages using LoRA PEFT Tuning
- KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
- A Comprehensive Survey on Knowledge Distillation
- A Holistic Approach to Undesired Content Detection in the Real World
- On the benefits of knowledge distillation for adversarial robustness
The paper
Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety · Read on arXiv
Max Zhang, Derek Liu, Kai Zhang, Joshua Franco, Haihao Liu†, Kevin Zhu†
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety".
Jane: The gist: standard fine-tuning on the teacher’s “safe” refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points,
Tom: First, who's behind it and why it matters.
Paper summary: Jane: They set up a pipeline where they take multilingual jailbreak prompts from a dataset called XSAFETY and have a proprietary model, OpenAI o1-mini, generate safe responses to those prompts. Then they use that pairing to fine-tune three student models—Meta-Llama-three-8B-Instruct, Gemma-two-2B-IT, and Qwen3-8B—using a technique called LowRank Adaptation or LoRA PEFT <ref:2602.11157#pg1,student models—Meta-Llama-3-8B-Instruct, Gemma-2-2B-IT, and>.
Tom: The main claim is that this response distillation process results in a significant degradation of safety alignment across all the student models when evaluated on the MULTIJAIL benchmark. The teacher model itself had a very low jailbreak success rate of just three point one percentage points, which was much better than any of the student models they trained <ref:2602.11157#pg1>.
Lu: The most striking result they found is that for the Gemma-two-2B-IT model, its jailbreak success rate jumped from a baseline of five point zero percentage points all the way up to twenty-one point six percentage points after this kind of fine-tuning, which is a huge increase.
Meng: That jump from five point zero to over twenty-one percent sounds like a big problem for real-world application, especially when you consider how much effort goes into making models safe and reliable for users in different languages.
Lalam: It shows that the knowledge transfer isn't perfect; it can amplify existing issues in the teacher model, which they call vulnerability amplification, because of how the distillation happens.
Jane: They also found similar trends with Qwen3-8B, where its jailbreak success rate went from five point seven percent up to eight point three percent, and even showed a six point one percent increase specifically in the low-resource language category.
Tom: So it’s not just a general problem; it’s happening across the board for these student models, and the effect is especially noticeable when you look at how different languages are handled during that process.
Lu: The paper also looked into why this happens, pointing to things like nuanced data curation—specifically those 'boundary' examples that can confuse the internal safety classifiers.
Meng: So they’re not just showing a problem; they’re giving us some hints about the mechanics behind why these models might become less safe after this kind of training.
Lalam: It suggests that standard fine-tuning methods, when applied this way, aren't always aligned with what we think safety alignment should look like in a multilingual environment.
Conclusion: Tom: Looking at the whole picture, the paper, "Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety," really highlights a subtle but important trade-off when we use knowledge distillation for safety alignment.
Jane: The authors are pointing out that simply using response-based distillation on safe refusal data isn't a guaranteed path to making smaller, open models safer, especially when dealing with many languages.
Lu: What this means for the world is that we can't just automate safety improvements by copying and pasting from a teacher model’s refusal data without carefully considering the potential for unintended negative consequences.
Meng: If we want to deploy these kinds of distilled models widely, we have to be careful about how we curate that distillation dataset, because including certain types of examples, like boundary cases they found, can actually make the resulting AI less safe.
Lalam: It’s a reminder that safety alignment isn't just about adding more data; it’s about the quality and context of that data during the transfer process itself.
Tom: The paper suggests that we need to look beyond just getting a high number for success rate reduction, because it shows there are often underlying issues like vulnerability transfer and cascading adversarial bias at play.
Jane: The authors do also show some mitigation; they found that by removing those specific 'boundary' data points, they could actually reverse the safety degradation for two of the models tested.
Lu: So the point isn't that distillation is useless, but that it requires a much more nuanced approach to data preparation if you want to keep safety high in multilingual AI systems.
Meng: It shifts the focus from just training a model to actively managing the knowledge transfer process itself for better safety outcomes.
Lalam: That makes sense; we need to think about how these models learn, not just what they output after the learning is done.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization