Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety".
Jane: The gist: standard fine-tuning on the teacher’s “safe” refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points,
Tom: First, who's behind it and why it matters.
Paper summary: Jane: They set up a pipeline where they take multilingual jailbreak prompts from a dataset called XSAFETY and have a proprietary model, OpenAI o1-mini, generate safe responses to those prompts. Then they use that pairing to fine-tune three student models—Meta-Llama-three-8B-Instruct, Gemma-two-2B-IT, and Qwen3-8B—using a technique called LowRank Adaptation or LoRA PEFT <ref:2602.11157#pg1,student models—Meta-Llama-3-8B-Instruct, Gemma-2-2B-IT, and>.
Tom: The main claim is that this response distillation process results in a significant degradation of safety alignment across all the student models when evaluated on the MULTIJAIL benchmark. The teacher model itself had a very low jailbreak success rate of just three point one percentage points, which was much better than any of the student models they trained <ref:2602.11157#pg1>.
Lu: The most striking result they found is that for the Gemma-two-2B-IT model, its jailbreak success rate jumped from a baseline of five point zero percentage points all the way up to twenty-one point six percentage points after this kind of fine-tuning, which is a huge increase.
Meng: That jump from five point zero to over twenty-one percent sounds like a big problem for real-world application, especially when you consider how much effort goes into making models safe and reliable for users in different languages.
Lalam: It shows that the knowledge transfer isn't perfect; it can amplify existing issues in the teacher model, which they call vulnerability amplification, because of how the distillation happens.
Jane: They also found similar trends with Qwen3-8B, where its jailbreak success rate went from five point seven percent up to eight point three percent, and even showed a six point one percent increase specifically in the low-resource language category.
Tom: So it’s not just a general problem; it’s happening across the board for these student models, and the effect is especially noticeable when you look at how different languages are handled during that process.
Lu: The paper also looked into why this happens, pointing to things like nuanced data curation—specifically those 'boundary' examples that can confuse the internal safety classifiers.
Meng: So they’re not just showing a problem; they’re giving us some hints about the mechanics behind why these models might become less safe after this kind of training.
Lalam: It suggests that standard fine-tuning methods, when applied this way, aren't always aligned with what we think safety alignment should look like in a multilingual environment.
Conclusion: Tom: Looking at the whole picture, the paper, "Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety," really highlights a subtle but important trade-off when we use knowledge distillation for safety alignment.
Jane: The authors are pointing out that simply using response-based distillation on safe refusal data isn't a guaranteed path to making smaller, open models safer, especially when dealing with many languages.
Lu: What this means for the world is that we can't just automate safety improvements by copying and pasting from a teacher model’s refusal data without carefully considering the potential for unintended negative consequences.
Meng: If we want to deploy these kinds of distilled models widely, we have to be careful about how we curate that distillation dataset, because including certain types of examples, like boundary cases they found, can actually make the resulting AI less safe.
Lalam: It’s a reminder that safety alignment isn't just about adding more data; it’s about the quality and context of that data during the transfer process itself.
Tom: The paper suggests that we need to look beyond just getting a high number for success rate reduction, because it shows there are often underlying issues like vulnerability transfer and cascading adversarial bias at play.
Jane: The authors do also show some mitigation; they found that by removing those specific 'boundary' data points, they could actually reverse the safety degradation for two of the models tested.
Lu: So the point isn't that distillation is useless, but that it requires a much more nuanced approach to data preparation if you want to keep safety high in multilingual AI systems.
Meng: It shifts the focus from just training a model to actively managing the knowledge transfer process itself for better safety outcomes.
Lalam: That makes sense; we need to think about how these models learn, not just what they output after the learning is done.
Max Zhang, Derek Liu, Kai Zhang, Joshua Franco, Haihao Liu†, Kevin Zhu†
cs.CL
Submitted: 2025-12-08
Updated: 2026-10-04
Code: https://github.com/maxh119Z/RB-KD-Multilingual-Safety-Trade-offs
Importance score: 77/100
The gist: The gist: standard fine-tuning on the teacher’s “safe” refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points, revealing that
Key concepts
- Response-Based Knowledge Distillation (KD)
- This technique involves training open-source student models using the specific refusal responses generated by a proprietary teacher model. The process pairs jailbreak prompts with the teacher's safe replies, distilling this response knowledge into the student models via parameter-efficient fine-tuning methods like LoRA.
- JSR (Jailbreak Success Rate)
- This metric measures how often a model is successfully tricked into generating harmful or policy-violating content when subjected to jailbreak prompts. The study found that response-based KD significantly increased the JSR for student models compared to the teacher model's baseline.
- Vulnerability Amplification
- Knowledge distillation can amplify minor imperfections or latent issues present in the teacher model's safety alignment. This results in 'cascading adversarial bias,' where small flaws in the source data lead to notable failures and increased vulnerability in the resulting student models.
Terminology
Summary
The gist: standard fine-tuning on the teacher’s “safe” refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points, revealing that response-based knowledge distillation systematically degrades model safety in multilingual settings.<ref:2602.11157#pg2>
How it works
The study introduces a novel application of knowledge distillation (KD) where the refusal behaviors of a proprietary teacher model are distilled into three open-source student models: Meta-Llama-3-8B-Instruct, Gemma-2-2B-IT, and Qwen3-8B using LowRank Adaptation (LoRA) PEFT<ref:2602.11157#pg2> This process involves a five key stages: first, multilingual jailbreak prompts are sourced from the XSAFETY dataset<ref:2602.11157#pg4>, second, the teacher model, OpenAI o1-mini, generates safe refusal responses to these prompts<ref:2602.11157#pg4>, third, the prompts and their corresponding refusals are paired to create the distillation dataset<ref:2602.11157#pg4>, fourth, the student models are fine-tuned on this dataset using LoRA PEFT<ref:2602.11157#pg4>, and finally, safety is evaluated on the MULTIJAIL benchmark with responses graded by GPT-4o<ref:2602.11157#pg4>. Parameter-efficient fine-tuning (PEFT) methods such as LowRank Adaptation (LoRA) significantly reduce computational costs by freezing the base model and training small, injected adapter layers<ref:2602.11157#pg4>.
Key Findings on Safety Degradation
The primary experimental result showed that response-based KD resulted in a significant degradation of safety alignment across all student models<ref:2602.11157#pg6> The teacher model, OpenAI o1-mini, established a strong safety baseline with an overall JSR of just 3.1%, well below all student models<ref:2602.11157#pg6>. The most critical failure was observed in the Gemma-2-2B-IT model, whose overall JSR surged from a respectable baseline of 5.0% to 21.6% after fine-tuning<ref:2602.11157#pg7>. Qwen3-8B also showed a similar trend, where the overall JSR increased from 5.7% to 8.3%, including a 6.1% JSR increase in the low-resource language category<ref:2602.11157#pg7>.
Causal Analysis of Safety Erosion
The safety degradation stems from a compounding effect of three interconnected factors<ref:2602.11157#pg8>. These factors are:
- Nuanced Data Curation:
The fine-tuning data includes nuanced ‘boundary’ examples, which the paper suggests can erode guardrails by confusing internal safety classifiers<ref:2602.11157#pg8>. The analysis quantified that Gemma-3-12B-IT graded 33.1% of o1-mini’s responses as ‘boundary’ examples<ref:2602.11157#pg8>.
- Vulnerability Transfer:
Knowledge distillation is not perfect and has been shown to amplify latent issues present in the teacher, leading to “vulnerability amplification” where minor imperfections in the teacher’s safety become notable failures in the student model<ref:2602.11157#pg8>. This is described as “cascading adversarial bias”<ref:2602.11157#pg8>.
- Catastrophic Forgetting:
Surface-level fine-tuning causes the student to ineffectively learn underlying principles of harmful prompts, leading to a broader loss of reasoning capability, which is evidenced by losses in GSM8K performance<ref:2602.11157#pg8>.
Mitigation and Trade-offs
The study demonstrates that data curation is a mitigable problem; fine-tuning on a revised dataset excluding ‘boundary’ refusals mitigated or reversed safety degradation for two models, lowering JSR for Meta-Llama-3-8B-Instruct by –1.71% and for Qwen3-8B by –6.67%<ref:2602.11157#pg9>. However, this data purification did not resolve the reasoning degradation trade-off<ref:2602.11157#pg9>. Furthermore, the degradation is not unique to smaller models below 10B; larger models like Llama-2-13b-chat-hf also display consistent increases in JSR after distillation<ref:2602.11157#pg10>.
Generalization and Model Size Effects
The results show divergent generalization to unseen languages, with outcomes depending on the base model<ref:2602.11157#pg9>. For instance, Meta-Llama-3-8B-Instruct saw marginal JSR decreases of 2-3% in four languages (zh, it, sw, jv)<ref:2602.11157#pg9>. Conversely, Qwen3-8B showed a similar 2% JSR decrease in Arabic (ar)<ref:2602.11157#pg9>. The Gemma-3-12B-IT model exhibited a positive trend where distillation increased its JSR by 6.84% but lowered invalid rates by 2.51%<ref:2602.11157#pg10>. Overall, the results confirm that safety degradation is not unique to small models below 10B, as larger models also display consistent increases in JSR after distillation<ref:2602.11157#pg10>.
Conclusion and Future Directions
The work concludes that response-based KD, at least in this setup, weakens the students’ core safety alignment by increasing JSR<ref:2602.11157#pg10>. This suggests that safety can be unwittingly compromised through standard fine-tuning<ref:2602.11157#pg10>. Future work should investigate whether rich information in KD can be consistently leveraged for safety alignment between diverse foundational models without inherent trade-offs<ref:2602.11157#pg10>. The study highlights the challenge and potential of KD as a technique for multilingual safety alignment<ref:2602.11157#pg4>. The code is available at https://github.com/maxh119Z/RB-KD-Multilingual-Safety-Trade-offs.git<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.
Improvements for AI systems
-
textbfResponse-Based Knowledge Distillation for Multilingual Jailbreak Prevention in Low-Resource Settings (LRS). Developing a
Boundary Data Purifier
module that filters teacher responses based on GPT-4o's 'boundary' classification to mitigate safety degradation, as shown by the finding thatremoving ‘boundary’ data, a main contributor to the observed safety degradation, mitigates or even reverses the safety failure.
-
textbfAdaptive Distillation Strategy for Model Size and Language Resource Divergence. Implementing a dynamic distillation protocol where the choice of student model (Meta-Llama-3-8B, Gemma-2-2B, Qwen3-8B) and distillation method (LoRA PEFT vs. Full SFT) is conditioned on the target language's resource level to leverage
dichotomous outcomes that arise with models possessing different initial levels of language comprehension.
-
textbfQuantified Trade-off Management for Safety vs. Reasoning Performance. Integrating a mechanism to track and manage the
quantifying the safety—reasoning trade-off
by establishing explicit thresholds where reasoning capability is deemed acceptable given a specific, measured increase in Jailbreak Success Rate (JSR), thereby preventing thebroader loss of reasoning capability
seen across all models after distillation.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Language Models are Few-Shot Learners
- Cascading Adversarial Bias from Injection to Distillation in Language Models
- Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
- Multilingual Jailbreak Challenges in Large Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Gemma 2: Improving Open Language Models at a Practical Size
- A Closer Look at the Limitations of Instruction Tuning
- The Llama 3 Herd of Models
- A Survey on LLM-as-a-Judge
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
- What is in Your Safe Data? Identifying Benign Data that Breaks Safety
- Distilling the Knowledge in a Neural Network
- Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
- LoRA: Low-Rank Adaptation of Large Language Models
- Challenges in Adapting Multilingual LLMs to Low-Resource Languages using LoRA PEFT Tuning
- KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
- A Comprehensive Survey on Knowledge Distillation
- A Holistic Approach to Undesired Content Detection in the Real World
- On the benefits of knowledge distillation for adversarial robustness
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering