Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety

arXiv:2602.11157 · cs.CL · Submitted 2025-12-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety".

Jane: The gist: standard fine-tuning on the teacher’s “safe” refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points,

Tom: First, who's behind it and why it matters.

Paper summary: Jane: They set up a pipeline where they take multilingual jailbreak prompts from a dataset called XSAFETY and have a proprietary model, OpenAI o1-mini, generate safe responses to those prompts. Then they use that pairing to fine-tune three student models—Meta-Llama-three-8B-Instruct, Gemma-two-2B-IT, and Qwen3-8B—using a technique called LowRank Adaptation or LoRA PEFT <ref:2602.11157#pg1,student models—Meta-Llama-3-8B-Instruct, Gemma-2-2B-IT, and>.

Tom: The main claim is that this response distillation process results in a significant degradation of safety alignment across all the student models when evaluated on the MULTIJAIL benchmark. The teacher model itself had a very low jailbreak success rate of just three point one percentage points, which was much better than any of the student models they trained <ref:2602.11157#pg1>.

Lu: The most striking result they found is that for the Gemma-two-2B-IT model, its jailbreak success rate jumped from a baseline of five point zero percentage points all the way up to twenty-one point six percentage points after this kind of fine-tuning, which is a huge increase.

Meng: That jump from five point zero to over twenty-one percent sounds like a big problem for real-world application, especially when you consider how much effort goes into making models safe and reliable for users in different languages.

Lalam: It shows that the knowledge transfer isn't perfect; it can amplify existing issues in the teacher model, which they call vulnerability amplification, because of how the distillation happens.

Jane: They also found similar trends with Qwen3-8B, where its jailbreak success rate went from five point seven percent up to eight point three percent, and even showed a six point one percent increase specifically in the low-resource language category.

Tom: So it’s not just a general problem; it’s happening across the board for these student models, and the effect is especially noticeable when you look at how different languages are handled during that process.

Lu: The paper also looked into why this happens, pointing to things like nuanced data curation—specifically those 'boundary' examples that can confuse the internal safety classifiers.

Meng: So they’re not just showing a problem; they’re giving us some hints about the mechanics behind why these models might become less safe after this kind of training.

Lalam: It suggests that standard fine-tuning methods, when applied this way, aren't always aligned with what we think safety alignment should look like in a multilingual environment.

Conclusion: Tom: Looking at the whole picture, the paper, "Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety," really highlights a subtle but important trade-off when we use knowledge distillation for safety alignment.

Jane: The authors are pointing out that simply using response-based distillation on safe refusal data isn't a guaranteed path to making smaller, open models safer, especially when dealing with many languages.

Lu: What this means for the world is that we can't just automate safety improvements by copying and pasting from a teacher model’s refusal data without carefully considering the potential for unintended negative consequences.

Meng: If we want to deploy these kinds of distilled models widely, we have to be careful about how we curate that distillation dataset, because including certain types of examples, like boundary cases they found, can actually make the resulting AI less safe.

Lalam: It’s a reminder that safety alignment isn't just about adding more data; it’s about the quality and context of that data during the transfer process itself.

Tom: The paper suggests that we need to look beyond just getting a high number for success rate reduction, because it shows there are often underlying issues like vulnerability transfer and cascading adversarial bias at play.

Jane: The authors do also show some mitigation; they found that by removing those specific 'boundary' data points, they could actually reverse the safety degradation for two of the models tested.

Lu: So the point isn't that distillation is useless, but that it requires a much more nuanced approach to data preparation if you want to keep safety high in multilingual AI systems.

Meng: It shifts the focus from just training a model to actively managing the knowledge transfer process itself for better safety outcomes.

Lalam: That makes sense; we need to think about how these models learn, not just what they output after the learning is done.

Max Zhang, Derek Liu, Kai Zhang, Joshua Franco, Haihao Liu†, Kevin Zhu†

cs.CL

Submitted: 2025-12-08

Updated: 2026-10-04

Code: https://github.com/maxh119Z/RB-KD-Multilingual-Safety-Trade-offs

Importance score: 77/100

The gist: The gist: standard fine-tuning on the teacher’s “safe” refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points, revealing that

Key concepts

Response-Based Knowledge Distillation (KD)
This technique involves training open-source student models using the specific refusal responses generated by a proprietary teacher model. The process pairs jailbreak prompts with the teacher's safe replies, distilling this response knowledge into the student models via parameter-efficient fine-tuning methods like LoRA.
JSR (Jailbreak Success Rate)
This metric measures how often a model is successfully tricked into generating harmful or policy-violating content when subjected to jailbreak prompts. The study found that response-based KD significantly increased the JSR for student models compared to the teacher model's baseline.
Vulnerability Amplification
Knowledge distillation can amplify minor imperfections or latent issues present in the teacher model's safety alignment. This results in 'cascading adversarial bias,' where small flaws in the source data lead to notable failures and increased vulnerability in the resulting student models.

Terminology

Summary

The gist: standard fine-tuning on the teacher’s “safe” refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points, revealing that response-based knowledge distillation systematically degrades model safety in multilingual settings.<ref:2602.11157#pg2>

How it works

The study introduces a novel application of knowledge distillation (KD) where the refusal behaviors of a proprietary teacher model are distilled into three open-source student models: Meta-Llama-3-8B-Instruct, Gemma-2-2B-IT, and Qwen3-8B using LowRank Adaptation (LoRA) PEFT<ref:2602.11157#pg2> This process involves a five key stages: first, multilingual jailbreak prompts are sourced from the XSAFETY dataset<ref:2602.11157#pg4>, second, the teacher model, OpenAI o1-mini, generates safe refusal responses to these prompts<ref:2602.11157#pg4>, third, the prompts and their corresponding refusals are paired to create the distillation dataset<ref:2602.11157#pg4>, fourth, the student models are fine-tuned on this dataset using LoRA PEFT<ref:2602.11157#pg4>, and finally, safety is evaluated on the MULTIJAIL benchmark with responses graded by GPT-4o<ref:2602.11157#pg4>. Parameter-efficient fine-tuning (PEFT) methods such as LowRank Adaptation (LoRA) significantly reduce computational costs by freezing the base model and training small, injected adapter layers<ref:2602.11157#pg4>.

Key Findings on Safety Degradation

The primary experimental result showed that response-based KD resulted in a significant degradation of safety alignment across all student models<ref:2602.11157#pg6> The teacher model, OpenAI o1-mini, established a strong safety baseline with an overall JSR of just 3.1%, well below all student models<ref:2602.11157#pg6>. The most critical failure was observed in the Gemma-2-2B-IT model, whose overall JSR surged from a respectable baseline of 5.0% to 21.6% after fine-tuning<ref:2602.11157#pg7>. Qwen3-8B also showed a similar trend, where the overall JSR increased from 5.7% to 8.3%, including a 6.1% JSR increase in the low-resource language category<ref:2602.11157#pg7>.

Causal Analysis of Safety Erosion

The safety degradation stems from a compounding effect of three interconnected factors<ref:2602.11157#pg8>. These factors are:

- Nuanced Data Curation:

The fine-tuning data includes nuanced ‘boundary’ examples, which the paper suggests can erode guardrails by confusing internal safety classifiers<ref:2602.11157#pg8>. The analysis quantified that Gemma-3-12B-IT graded 33.1% of o1-mini’s responses as ‘boundary’ examples<ref:2602.11157#pg8>.

- Vulnerability Transfer:

Knowledge distillation is not perfect and has been shown to amplify latent issues present in the teacher, leading to “vulnerability amplification” where minor imperfections in the teacher’s safety become notable failures in the student model<ref:2602.11157#pg8>. This is described as “cascading adversarial bias”<ref:2602.11157#pg8>.

- Catastrophic Forgetting:

Surface-level fine-tuning causes the student to ineffectively learn underlying principles of harmful prompts, leading to a broader loss of reasoning capability, which is evidenced by losses in GSM8K performance<ref:2602.11157#pg8>.

Mitigation and Trade-offs

The study demonstrates that data curation is a mitigable problem; fine-tuning on a revised dataset excluding ‘boundary’ refusals mitigated or reversed safety degradation for two models, lowering JSR for Meta-Llama-3-8B-Instruct by –1.71% and for Qwen3-8B by –6.67%<ref:2602.11157#pg9>. However, this data purification did not resolve the reasoning degradation trade-off<ref:2602.11157#pg9>. Furthermore, the degradation is not unique to smaller models below 10B; larger models like Llama-2-13b-chat-hf also display consistent increases in JSR after distillation<ref:2602.11157#pg10>.

Generalization and Model Size Effects

The results show divergent generalization to unseen languages, with outcomes depending on the base model<ref:2602.11157#pg9>. For instance, Meta-Llama-3-8B-Instruct saw marginal JSR decreases of 2-3% in four languages (zh, it, sw, jv)<ref:2602.11157#pg9>. Conversely, Qwen3-8B showed a similar 2% JSR decrease in Arabic (ar)<ref:2602.11157#pg9>. The Gemma-3-12B-IT model exhibited a positive trend where distillation increased its JSR by 6.84% but lowered invalid rates by 2.51%<ref:2602.11157#pg10>. Overall, the results confirm that safety degradation is not unique to small models below 10B, as larger models also display consistent increases in JSR after distillation<ref:2602.11157#pg10>.

Conclusion and Future Directions

The work concludes that response-based KD, at least in this setup, weakens the students’ core safety alignment by increasing JSR<ref:2602.11157#pg10>. This suggests that safety can be unwittingly compromised through standard fine-tuning<ref:2602.11157#pg10>. Future work should investigate whether rich information in KD can be consistently leveraged for safety alignment between diverse foundational models without inherent trade-offs<ref:2602.11157#pg10>. The study highlights the challenge and potential of KD as a technique for multilingual safety alignment<ref:2602.11157#pg4>. The code is available at https://github.com/maxh119Z/RB-KD-Multilingual-Safety-Trade-offs.git<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.11157#pg4>. The paper is presented and published at NeurIPS 2025 ResponsibleFM Workshop<ref:2602.11157#pg4>. The arXiv identifier is 2602.11157v1[cs.CL] dated 8 Dec 2025<ref:2602.

Improvements for AI systems

  1. textbfResponse-Based Knowledge Distillation for Multilingual Jailbreak Prevention in Low-Resource Settings (LRS). Developing a Boundary Data Purifier module that filters teacher responses based on GPT-4o's 'boundary' classification to mitigate safety degradation, as shown by the finding that removing ‘boundary’ data, a main contributor to the observed safety degradation, mitigates or even reverses the safety failure.

  2. textbfAdaptive Distillation Strategy for Model Size and Language Resource Divergence. Implementing a dynamic distillation protocol where the choice of student model (Meta-Llama-3-8B, Gemma-2-2B, Qwen3-8B) and distillation method (LoRA PEFT vs. Full SFT) is conditioned on the target language's resource level to leverage dichotomous outcomes that arise with models possessing different initial levels of language comprehension.

  3. textbfQuantified Trade-off Management for Safety vs. Reasoning Performance. Integrating a mechanism to track and manage the quantifying the safety—reasoning trade-off by establishing explicit thresholds where reasoning capability is deemed acceptable given a specific, measured increase in Jailbreak Success Rate (JSR), thereby preventing the broader loss of reasoning capability seen across all models after distillation.

Sources

Related papers