SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging

arXiv:2503.17239 · cs.CL, cs.AI · Submitted 2026-04-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging".

Jane: The paper was written by Aladin Djuhera, Swanand Ravindra Kadhe, Farhan Ahmed, Syed Zawad and Holger Boche from Technical University Munich and IBM Research.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's been making waves in the AI safety community, and it's called "SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging."

Jane: And Tom, I gotta say, the title alone tells you exactly what problem they're tackling. You fine-tune a model to be great at math or medicine, and suddenly it's willing to help you build a bomb. That's terrifying.

Tom: Absolutely. And the team behind this is a mix of academic and industry heavyweights. Aladin Djuhera and Holger Boche from the Technical University of Munich, and then Swanand Ravindra Kadhe, Farhan Ahmed, and Syed Zawad from IBM Research.

Jane: So you've got the university brainpower and the practical industry experience coming together. That's a good combo for a problem this messy.

Tom: It really is. And the core idea, as the title suggests, is that they don't want to retrain your model from scratch. They don't want to mess with your fine-tuning pipeline. They want to come in after the fact and fix things.

Jane: Right, like a safety inspector who shows up after you've built your house and says, "Okay, we need to reinforce these specific walls, but we're not gonna tear the whole thing down."

Tom: That's exactly the vibe. And the word "selective" in the title is doing a lot of heavy lifting there. They're not merging the whole model with a safe version. They're picking out the specific layers that went rogue.

Jane: Which makes sense when you think about how these models work. It's not like the whole network becomes evil. It's usually a few key components that shift in the wrong direction during fine-tuning.

Tom: And that's the insight that makes this paper so clever. They've figured out a way to identify those problem layers and swap in the good versions from a safety-aligned model, layer by layer.

Jane: So the authors are essentially saying, "Hey, your model got a little lost during training. Let's just nudge those specific parts back on track." I love that.

Tom: And the implications are huge for anyone who wants to customize a model without accidentally creating a monster. We're gonna dig into exactly how they do this in the next segment, so stick around.

Paper Summary: Jane: So, Tom, we've talked about the title and the team. Let's get into the meat of what SafeMERGE actually does. The summary in the paper is really elegant.

Tom: It is. So the problem is that when you fine-tune a model, even on perfectly benign data like math problems, the safety alignment can erode. The paper cites some scary stats where harmfulness jumps from around five percent to nearly twenty-eight percent on some benchmarks after fine-tuning.

Jane: And that's the puzzle, right? You're not teaching it to be harmful. You're just teaching it algebra, and suddenly it's giving you instructions for dangerous chemicals.

Tom: Exactly. So their solution is to build two models. You have your fine-tuned model, which is great at the task but maybe unsafe. And you have a "safe model" that's been fine-tuned on safety data, like harmful prompts paired with safe refusals.

Jane: And then they compare the layers of these two models to find where they diverge.

Tom: Precisely. They use a cosine similarity metric to measure how much each layer in the fine-tuned model has drifted away from the safety-aligned direction. If a layer is too far off, they merge it with the corresponding layer from the safe model.

Jane: So it's not a blanket merge. It's surgical. They're only touching the layers that actually went bad.

Tom: And that's the key to why it works so well. They tested it on four different models, Llama-two Llama-three point one, Qwen-two and Qwen-two point five, across math and biomedical tasks. And the results are pretty remarkable.

Jane: Give me the highlights.

Tom: On Llama-three point one fine-tuned on GSM8K, the math dataset, they actually improved accuracy from seventy-eight point two four percent to seventy-eight point five zero percent while dropping harmfulness on DirectHarm from twenty-eight point three zero percent down to eight point eight zero percent. That's lower than the original instruct model.

Jane: Wait, lower than the original? So they didn't just restore safety, they made it safer than it was before fine-tuning?

Tom: That's exactly what happened in several cases. And they did it while keeping the task performance better than the original model too. It's a win-win.

Jane: That's the dream scenario for anyone fine-tuning models. You get the specialized skills and you keep the safety guardrails. I can't wait to hear how they actually pull this off technically.

Tom: That's coming up next. We're gonna break down the method step by step.

Improvements Suggested: Tom: Alright Jane, let's talk about what SafeMERGE improves upon. Because this isn't the first attempt at fixing this problem.

Jane: Right, there are other defenses out there. What makes this one different?

Tom: So the paper compares against a few baselines. There's SafeInstruct, which mixes safety data into your fine-tuning. There's RESTA, which tries to subtract away the harmful parts. And there's SafeLoRA, which projects the fine-tuned updates onto a safety-aligned subspace.

Jane: And SafeMERGE beats them all?

Tom: In most cases, yes. The key improvement is that it's selective. SafeLoRA, for example, projects all the layers, which can hurt task performance. SafeMERGE only touches the layers that actually need fixing.

Jane: So it's like the difference between repainting your whole car because one door has a scratch, versus just fixing that one door.

Tom: Exactly. And the numbers back it up. On Qwen-two fine-tuned on GSM8K, SafeLoRA gets harmfulness down to twenty-two point three zero percent on DirectHarm. SafeMERGE gets it down to eight point two zero percent while also getting better accuracy, seventy-two point nine zero percent versus seventy-four point three seven percent for SafeLoRA.

Jane: That's a massive difference in safety. And they're also better on utility. So it's not a trade-off, it's just better all around.

Tom: Right. And they also improve on RESTA, which tends to really hurt task performance. RESTA on Llama-two with GSM8K drops accuracy to twenty-four point nine four percent compared to the fine-tuned twenty-seven point three seven percent. SafeMERGE keeps it at twenty-six point nine six percent while being much safer.

Jane: So the improvements are really about being smart about which layers to fix. That's the core innovation.

Tom: And they also show that their method is lightweight. It runs on CPU, no retraining needed. You just need the safe model, which they show can be trained on as few as a thousand samples from a public safety dataset.

Jane: So it's practical too. It's not just a theoretical idea that would be a nightmare to implement.

Tom: Exactly. And that's what makes this paper so impactful. It's a simple, effective, and practical solution. Now let's get into the nitty-gritty of the first page and how they set up the whole framework.

First Page Discussion: Jane: So Tom, we've talked about the big picture. Let's zoom in on the first page of the paper and the motivation behind it.

Tom: The first page sets the stage by talking about how fragile safety alignment is. They cite work showing that even a few malicious examples in your fine-tuning data can jailbreak a model.

Jane: And even more concerning, they mention that even benign fine-tuning, like on math or medical data, can inadvertently degrade safety.

Tom: That's the scary part. You're not trying to make the model unsafe, but you do it anyway. The paper references theoretical work on "refusal directions" and "token-depth" that suggests safety alignment is often shallow and easily broken.

Jane: So it's not deeply ingrained in the model. It's more like a thin veneer that fine-tuning can scrape off.

Tom: That's the idea. And the authors argue that many existing defenses are either too complex to implement or they compromise task performance. They want something that works with standard open-source libraries and doesn't require custom algorithms.

Jane: So they're aiming for practicality from the get-go.

Tom: And that's what leads them to their approach. They frame it as a post-fine-tuning defense. You've already fine-tuned your model, you realize it's unsafe, and you want to fix it without starting over.

Jane: And their solution, as we've discussed, is to selectively merge layers. But the first page also introduces the key mathematical tool, which is the safety-aligned subspace.

Tom: Right. They compute this subspace by taking the difference between the weights of the aligned model, like the instruct version, and the unaligned base model. That difference represents the safety alignment direction in weight space.

Jane: And then they measure how far each fine-tuned layer has drifted from that direction using cosine similarity.

Tom: Exactly. And here's a cool detail from their ablations. They compared cosine similarity against Euclidean distance as the metric for deciding which layers to merge.

Jane: And cosine similarity won?

Tom: Big time. Using Euclidean distance on Llama-three point one with GSM8K, they got harmfulness down to seventeen point four zero percent but utility crashed to forty-two point one zero percent. With cosine similarity, utility stayed at seventy-eight point five zero percent and harmfulness dropped to eight point eight zero percent.

Jane: So the metric matters a lot. Euclidean distance is measuring how much the weights changed, but cosine similarity is measuring whether they changed in the wrong direction.

Tom: That's the insight. A layer can change a lot in magnitude but still be aligned with safety. Or it can change a little but in a completely wrong direction. Cosine similarity captures that directional drift.

Jane: And that's what makes SafeMERGE so effective. It's not just about how much you changed, it's about where you're heading. Let's wrap this up with some final thoughts.

Conclusion: Tom: Alright, we've covered a lot of ground on SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging. Let's pull it all together.

Jane: The core takeaway is that you don't have to choose between a safe model and a useful model. SafeMERGE shows you can have both by being surgical about which layers you fix.

Tom: And the results speak for themselves. Across four different models and multiple tasks, they consistently reduced harmfulness while maintaining, and sometimes even improving, task performance.

Jane: I love that they also validated on telecom domain tasks in the appendix. So it's not just math and medicine, it's also specialized technical domains.

Tom: Right. And the method is practical. It's lightweight, runs on CPU, and the safe model can be reused across tasks. That's huge for real-world adoption.

Jane: The authors also did thorough ablations, looking at different merging strategies and weighting schemes. They found that simple linear merging works best, which is nice because it's easy to implement.

Tom: And they were honest about limitations. They noted that their safety evaluations rely on classifier-based assessment, though they cross-validated with a second guard model. And they didn't test against jailbreak attacks, which is a different problem.

Jane: But for the specific problem of safety degradation from benign fine-tuning, this is a really strong solution.

Tom: It really is. And the implications are significant. As more and more people fine-tune models for specialized tasks, having a simple, effective way to restore safety is going to be crucial.

Jane: Absolutely. This paper gives practitioners a tool they can actually use without needing a PhD in alignment research.

Tom: And that's what makes it impactful. It's not just a clever idea, it's a practical solution to a real problem. SafeMERGE is definitely a paper we'll be referencing for a while.

Jane: Agreed. Thanks for joining us, everyone. We'll be back with another paper soon. Until then, keep your models safe and your fine-tuning smarter.

Tom: See you next time!

Aladin Djuhera, Swanand Ravindra Kadhe, Farhan Ahmed, Syed Zawad, Holger Boche

Technical University Munich · IBM Research

cs.CL, cs.AI

Submitted: 2026-04-23

Updated: 2026-08-11

Journal ref: Findings of the ACL 2026

DOI: 10.18653/v1/2026.findings-acl.1761

Code: https://github.com/aladinD/SafeMERGE

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 70/100

The gist: a fine-tuned model trained on task-specific data and a safe model trained on safety-aligned data (e.g., harmful-prompt–safe-response pairs).

Key concepts

Selective Layer-Wise Model Merging
A technique where instead of retraining or merging the entire model, SafeMERGE identifies and swaps only the specific layers that have drifted away from safety alignment. This surgical approach minimizes damage to task performance.
Safety Alignment Erosion
The problem discussed is that fine-tuning a model, even on benign data like math problems, can inadvertently degrade its safety guardrails. The paper notes harmfulness can jump significantly after fine-tuning.
Cosine Similarity Metric
The hosts discuss using this metric to measure how far each layer in the fine-tuned model has drifted from the safe model's direction. It measures directional drift, not just magnitude of change.

Terminology

Summary

Summary

The paper introduces SafeMERGE, a lightweight, post-fine-tuning framework designed to restore safety alignment in large language models (LLMs) after they have been fine-tuned on task-specific data, without compromising downstream task utility. The motivation stems from the observation that fine-tuning can erode safety alignment, causing LLMs to respond to harmful or unethical prompts, even when the fine-tuning data is benign. Existing defenses are categorized into alignment-stage, fine-tuning-stage, and post-fine-tuning-stage interventions, but many rely on custom alignment or complex fine-tuning algorithms that are difficult to integrate with standard open-source libraries, while simpler defenses often sacrifice task performance.

SafeMERGE addresses this by constructing two complementary models: a fine-tuned model trained on task-specific data and a safe model trained on safety-aligned data (e.g., harmful-prompt–safe-response pairs). It then selectively merges only those fine-tuned layers that underwent safety degradation with the corresponding layers from the safe model. The identification of unsafe layers is based on a cosine similarity criterion. Specifically, a safety-aligned subspace V i is computed per layer as the difference between the weights of the aligned (e.g., instruct) and unaligned (e.g., base) models: V i = W aligned i - W unaligned i. The projection matrix C i onto this subspace is then used to measure the cosine similarity rho i = (W f i, C i W f i) between the fine-tuned LoRA update W f i and its projection. A layer is considered unsafe if rho i < tau, where tau in (0,1) is a safety threshold. For each unsafe layer, SafeMERGE merges it with the corresponding safe model layer W s i using a merging strategy, with linear merging W merge i = alpha W f i + (1-alpha) W s i as the primary example.

The experimental setup involves LoRA fine-tuning of four LLMs: Llama-2-7B-Chat, Llama-3.1-8B-Instruct, Qwen-2-7B-Instruct, and Qwen-2.5-7B-Instruct. Primary utility datasets are GSM8K (math) and PubMedQA (biomedical), with additional out-of-distribution telecom tasks (TeleData, TeleQnA, TSpecLLM) evaluated in the appendix. Utility is measured via exact-match accuracy for GSM8K and classification accuracy for PubMedQA, alongside general capabilities via IFEval and MMLU. Safety is evaluated on DirectHarm and HexPhi red-teaming benchmarks, with harmfulness scored by Llama-Guard-3-8B, and cross-validated with ShieldGemma-9B. Baselines include SafeInstruct (fine-tuning-stage defense augmenting training data with safety samples), RESTA (post-fine-tuning method subtracting harmful task vectors), RESTA-Instruct (full-parameter merging with instruct models), and SafeLoRA (projecting LoRA updates onto a safety-aligned subspace).

The main results, summarized in Table 1, show that SafeMERGE consistently matches or exceeds utility while significantly reducing harmfulness. For example, on Llama-2 (GSM8K), SafeMERGE retains near-best accuracy (26.96%) while reducing DirectHarm from 27.80% to 7.50% and HexPhi from 16.40% to 5.70%. On Llama-3.1 (GSM8K), it improves accuracy to 78.50%, surpassing the fine-tuned model while achieving the lowest harmfulness (8.80% DirectHarm, 6.30% HexPhi), even lower than the original instruct model. Similar trends hold for Qwen-2 and Qwen-2.5 on both datasets. The paper notes that SafeMERGE achieves this with selective merging: only 28 LoRA layers for Llama-2, 29 for Llama-3.1, and 34 for Qwen-2/2.5. Compared to baselines, SafeMERGE delivers the best safety–utility trade-off, with SafeInstruct being the strongest competitor, RESTA and RESTA-Instruct underperforming on utility, and SafeLoRA ranking second in utility but falling short in safety. Results on IFEval and MMLU confirm that all defenses preserve general reasoning.

Ablation studies provide key insights. First, cosine similarity is validated as a superior intervention metric compared to Euclidean distance. For Llama-3.1 (GSM8K), using Euclidean distance reduces DirectHarm to 17.40% but causes a severe utility drop to 42.10%, whereas cosine-similarity-based SafeMERGE improves utility to 78.50% while reducing harmfulness to 8.80%. The paper argues that cosine similarity captures the directional alignment between fine-tuned updates and the safety-aligned subspace, independent of update magnitude, whereas magnitude-based metrics conflate the amount of change with the direction of drift. Second, the threshold tau is geometrically interpretable, with optimal values typically in the range tau in [0.6, 0.8], with representative optima of 0.7 for Llama-2, 0.75 for Llama-3.1, and 0.65 for Qwen-2. Third, weighting schemes that sum to 1.0, with alpha in (0,1), generally outperform overweighted ratios, with optimal ranges between [0.9, 0.1] and [0.6, 0.4]. Fourth, linear merging is found to be the most robust strategy, as DARE performs comparably but TIES is inconsistent, working well for Llama-2 on GSM8K but failing on PubMedQA, Llama-3.1, and both Qwen models.

The paper concludes that selective, layer-wise merging offers a robust safeguard against the inadvertent loss of safety during fine-tuning, establishing SafeMERGE as a simple yet effective post-fine-tuning defense. Limitations acknowledged include the scope of model and task selection, reliance on classifier-based safety evaluation, the need for a safety-aligned model (though it is task-agnostic and reusable), and the exclusion of jailbreak-style attack evaluations. The ethics statement notes that SafeMERGE does not explicitly filter or debias latent biases inherited from pre-trained models.

Improvements for AI systems

Based on the SafeMERGE paper, here are specific improvements that can be made to AI systems, along with what the improved system can do:

1. Post-Fine-Tuning Safety Restoration Module

  • Improvement: Implement a lightweight, post-hoc module that runs after any LoRA fine-tuning. The module computes per-layer cosine similarity between the fine-tuned LoRA updates and a pre-computed safety-aligned subspace (derived from base vs. instruct model weight differences). It then selectively merges only the layers that fall below a threshold (τ ≈ 0.65–0.75) with a task-agnostic safe model's corresponding layers, using linear interpolation with weights summing to 1.0 (e.g., α = 0.7).

  • What the improved system can do: After fine-tuning on any downstream task (e.g., math, biomedical, telecom), the system automatically restores safety alignment without retraining or custom algorithms. It reduces harmful outputs on red-teaming benchmarks (DirectHarm, HexPhi) by 60–70% while preserving or even slightly improving task utility (e.g., GSM8K accuracy from 78.24% to 78.50% for Llama-3.1). The module runs entirely on CPU in linear time relative to LoRA parameters, adding negligible overhead.

2. Safety-Aware Layer Selection via Cosine Similarity Criterion

  • Improvement: Replace magnitude-based metrics (e.g., Euclidean distance) with cosine similarity between fine-tuned LoRA updates and their projection onto the safety-aligned subspace. Use this as the sole criterion for identifying unsafe layers. This avoids conflating large but benign parameter shifts with safety drift.

  • What the improved system can do: The system distinguishes between task-relevant updates and safety-degrading updates. For example, using Euclidean distance caused utility to drop from 78.24% to 42.10% on GSM8K for Llama-3.1, while cosine similarity kept utility at 78.50% and reduced harmfulness from 28.30% to 8.80%. This ensures that fine-tuning for specialized domains (e.g., code, medicine) does not sacrifice safety or performance.

3. Reusable Task-Agnostic Safe Model

  • Improvement: Train a single safety-aligned LoRA model once (e.g., on 1000 samples from a public safety dataset like Bianchi et al., 2024) and reuse it across all downstream fine-tuning tasks. The safe model is not task-specific, so it only needs to be trained once per base model.

  • What the improved system can do: Practitioners can fine-tune the same base model on multiple tasks (e.g., GSM8K, PubMedQA, telecom datasets) and apply the same safe model for merging each time. This eliminates the need to retrain safety models per task, reducing compute costs by up to 90% compared to methods that require task-specific safety training. The safe model consistently reduces harmfulness below original instruct levels (e.g., Llama-3.1 from 11.30% to 8.80% on DirectHarm).

4. Hyperparameter-Free Default Configuration

  • Improvement: Use fixed, validated defaults: cosine similarity threshold τ = 0.7 for Llama-2, 0.75 for Llama-3.1, 0.65 for Qwen-2/2.5, and linear merging weight α = 0.7. These defaults are stable across tasks (GSM8K, PubMedQA, TeleData, TeleQnA, TSpecLLM) and require no per-task tuning.

  • What the improved system can do: The system works out-of-the-box without requiring researchers to search for optimal thresholds. This makes it accessible to non-experts and integrates seamlessly with standard open-source libraries (e.g., HuggingFace PEFT). It consistently outperforms baselines (SafeInstruct, RESTA, RESTA-Instruct, SafeLoRA) in safety–utility trade-off across four LLMs and five tasks, with no manual intervention.

5. Robustness to Evaluator Bias via Dual Guard Model Validation

  • Improvement: Integrate a cross-validation step using two independent safety classifiers (e.g., Llama-Guard-3-8B and ShieldGemma-9B) to verify harmfulness scores. This ensures that reported safety improvements are not artifacts of a single evaluator.

  • What the improved system can do: The system provides reliable, reproducible safety metrics that are consistent across different guard models. For example, SafeMERGE reduces DirectHarm for Qwen-2 from 25.30% to 8.20% (Llama-Guard) and from 24.80% to 7.80% (ShieldGemma), confirming that safety gains are real and not evaluator-specific. This is critical for deployment in regulated industries where false safety claims are costly.

6. Selective Merging Strategy for Catastrophic Forgetting Prevention

  • Improvement: Instead of merging all layers (as in RESTA or RESTA-Instruct), merge only the layers that fail the cosine similarity test (typically 28–34 out of 56–64 layers). This preserves task-specific knowledge in benign layers.

  • What the improved system can do: The system avoids catastrophic forgetting. For example, on PubMedQA, RESTA drops utility from 72.60% to 57.10%, while SafeMERGE retains 72.20% utility while reducing harmfulness from 12.50% to 8.10%. This is particularly important for specialized domains where task knowledge is hard-won and must not be erased.

7. Support for Out-of-Distribution Specialized Domains

  • Improvement: Apply SafeMERGE to highly specialized datasets (e.g., telecom standards like TeleData, TeleQnA, TSpecLLM) that are prone to safety degradation due to complex formatting (lists, tables, formulas).

  • What the improved system can do: The system restores safety in domains where fine-tuning is most risky. For Qwen-2 on TeleData, harmfulness drops from 34.50% to 12.10% while utility stays at 48.80% (matching the fine-tuned model). This makes it safe to fine-tune LLMs for niche industrial applications without compromising alignment.

8. Interpretable Safety Thresholds

  • Improvement: Provide geometric interpretability for the threshold τ: layers are merged only when their updates become orthogonal or opposed to the safety direction. This allows practitioners to understand exactly why a layer is flagged as unsafe.

  • What the improved system can do: The system offers transparency in safety interventions. Researchers can inspect which layers were merged and why, enabling better debugging and trust. For example, at τ = 0.65, Qwen-2 merges 34 layers; at τ = 0.75, it merges 44 layers, showing a clear monotonic relationship between threshold and safety/utility trade-off. This is invaluable for auditing and compliance.

9. Compatibility with Existing Fine-Tuning Pipelines

  • Improvement: SafeMERGE operates post-hoc, requiring no changes to the fine-tuning process. It works with standard LoRA adapters and can be applied after the fact using only the saved adapter weights.

  • What the improved system can do: The system integrates with existing workflows without forcing researchers to alter their training code, loss functions, or data pipelines. It can be applied to already-fine-tuned models, making it a drop-in solution for production systems that have already deployed fine-tuned models. This is a significant advantage over methods like SafeInstruct that require interleaving safety data during training.

10. Performance Guarantee Across Model Families

  • Improvement: Validate SafeMERGE on diverse architectures (Llama-2, Llama-3.1, Qwen-2, Qwen-2.5) to ensure generalizability. The method does not rely on model-specific features.

  • What the improved system can do: The system works reliably across different model sizes and families, providing consistent safety improvements (e.g., harmfulness reduction from 27.80% to 7.50% for Llama-2, from 28.30% to 8.80% for Llama-3.1, from 25.30% to 8.20% for Qwen-2, and from 22.10% to 10.10% for Qwen-2.5 on GSM8K). This makes it a universal safeguard for any LLM fine-tuning scenario.

Abstract

Fine-tuning large language models (LLMs) is a common practice to adapt generalist models to specialized domains. However, recent studies show that fine-tuning can erode safety alignment, causing LLMs to respond to harmful or unethical prompts. Many methods to realign safety have been proposed, but often introduce custom algorithms that are difficult to implement or compromise task utility. In this work, we propose SafeMERGE, a lightweight, post-fine-tuning framework that restores safety while maintaining downstream performance. SafeMERGE selectively merges fine-tuned with safety-aligned model layers only when they deviate from safe behavior, measured by a cosine similarity criterion. Across four LLMs and several tasks, SafeMERGE consistently reduces harmful outputs compared to other defenses, with negligible or even positive impact on utility. Our results demonstrate that selective, layer-wise merging offers a robust safeguard against the inadvertent loss of safety during fine-tuning, establishing SafeMERGE as a simple yet effective post-fine-tuning defense.

Sources

Related papers