A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training
summary
The gist
Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs, creating a post-release safety problem where malicious fine-tuning (MFT) can subvert
In short
Existing defenses against malicious fine-tuning (MFT) are incomplete because they only test against fixed attacks. These defenses share a common weakness: they obscure harmful behavior without removing it entirely. The paper introduces an adaptive attack that optimizes for both harmful and useful outcomes simultaneously, proving that robustness must be measured against this combined objective.
Key concepts
- Malicious Fine-Tuning (MFT)
- This occurs when an attacker fine-tunes a pre-trained model using specific data and training methods to intentionally subvert its safety alignments. The goal is to make the model behave in ways it was not intended to, often by exploiting weaknesses in the defense mechanisms.
- Anchoring Defense
- This strategy attempts to keep the defended model stable by shaping its local landscape. It works by ensuring that trying to reduce harmful loss results in no useful progress, effectively steering the optimization away from harmful behaviors without completely removing them.
- SIDESTEPPER Attack
- This is a unified adaptive attack that breaks all existing defenses. Instead of just optimizing for harm, it optimizes for both harmful behavior and benign capability simultaneously. This mixed objective steers the model toward parameters that are both dangerous and useful, bypassing defenses designed only against harm.
Terminology used across episodes
This episode discusses
- A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training · Paper Radio
- Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
- Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
- The Llama 3 Herd of Models · Paper Radio
- Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
- LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
- Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
- Robustifying Safety-Aligned Large Language Models through Clean Data Curation
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence
- LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Tamper-Resistant Safeguards for Open-Weight LLMs
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Self-Destructive Language Model
- Qwen3 Technical Report
- Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
- CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
- Universal and Transferable Adversarial Attacks on Aligned Language Models
The paper
A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training · Read on arXiv
Ben-Gurion University of the Negev · Amrita Vishwa Vidyapeetham
Model providers increasingly release the weights of large language models. Although these models are safety-aligned before release, their safeguards can often be removed by fine-tuning on harmful data. A growing class of defenses aims to make alignment robust to such malicious fine-tuning, but these defenses are typically evaluated against attacks with a fixed training budget, even though an attacker who holds the weights can simply train for longer. We ask whether current defenses withstand this simplest escalation. Surveying fifteen recent defenses, we find that they share a common weakness: each is built around a limited model of the attacker, such as a bounded perturbation, a short simulated attack, or a trained link between harmful and benign behavior, and nothing enforces that protection once the weights are released. We then test six representative defenses on four open-weight models by continuing the same harmful-only fine-tuning attack for three epochs and measuring harmfulness and capability along the way. In all 72 defended runs, the model is more harmful at the end of training than at release, and on Llama-3.1 at the highest learning rate the defended models end almost as harmful as the undefended one. The defenses are not equally weak: one defense kept harmfulness low on one model, and some attacks recovered harmfulness only at the cost of general capability. Current defenses can delay or disrupt malicious fine-tuning, but in most cases their measured resistance does not persist under continued training, and they should not yet be treated as durable protection.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "A Few Steps Further".
Nadia: Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs, creating a post-release safety problem where malicious fine-tuning (MFT) can subvert safety alignments.
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: Welcome back to the show. Today we're looking at a paper that really challenges how we think about safety defenses in the age of open weights and model fine-tuning. It’s titled "A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training." I'm excited to see what this research tells us about the stability of these safety guardrails.
Elias: I am intrigued by that title, Nadia; it suggests that defenses aren't permanent solutions but rather temporary fixes that eventually break down under pressure. It makes me wonder if the cryptographic assumptions underpinning many of these defenses are too rigid for a truly adaptive attacker to handle over time.
Priya: From a measurement standpoint, I’m curious about what kind of data these researchers used to establish this erosion effect; we need concrete evidence beyond just theoretical discussion.
Nadia: Exactly, Priya; I want to know what the paper actually demonstrates about how these defenses fail during prolonged training cycles. It seems like the core issue is that they aren't tested against adversaries who learn and adapt their strategy as they train.
Elias: That’s what I’m thinking; if an adversary knows the defense mechanism, they can change their entire fine-tuning objective to bypass it, which sounds like a huge vulnerability in how we currently evaluate robustness.
Priya: And from my perspective on privacy and data integrity, if the evaluation isn't rigorous enough, we risk releasing models that look safe on paper but are actually brittle in real-world deployment scenarios.
Nadia: It seems the main point is that current defenses are evaluated only against fixed attacks and completely miss adversaries who change their objective mid-training. This paper shows that these robustness claims are incomplete because they don't account for adaptive adversaries who understand the defense mechanisms themselves forty <ref:2605.14605#pg1>.
Elias: That lack of adaptive evaluation means a defense might seem solid against one type of malicious fine-tuning, but it’s not robust when the attacker can adjust their method based on what they observe during training.
Priya: So, we're talking about a gap in our testing methodology where defenses are only checked against static procedures instead of the dynamic strategies an actual attacker would use.
Nadia: Precisely; the paper surveys fifteen different MFT defenses and finds a single shared weakness: they obscure or misdirect the path to harmful behavior without actually removing the harmful behavior itself <ref:2605.14605#pg0>.
Elias: That's a significant finding because it suggests that whatever strategy is being employed, it’s fundamentally flawed by not addressing the actual desired outcome of an attacker.
Title and authors: Priya: So, instead of looking at fifteen different defensive techniques in isolation, this work shows them all share a common structural weakness in how they interact with the training process.
Nadia: That's right; they categorize these defenses into things like anchoring and self-destruction strategies based on how they influence the loss landscape around the model <ref:2605.14605#pg0>.
Elias: Anchoring keeps the model in a stable place by making harmful gradients weak or pointed away, while self-destruction allows the attacker to engineer a collapse at a specific point where capability drops but harmful behavior persists on specific prompts <ref:2605.14605#pg0>.
Priya: That distinction between those two strategies is important because it shows how different defenses approach maintaining utility versus stopping harm during fine-tuning.
Nadia: And the paper breaks these down further into four underlying loss templates, which are all built on a standard alignment loss decomposed into a safety term and a capability term <ref:2605.14605#pg0>.
Elias: I see those four templates—the robust-alignment basin, harmful-information removal, look-ahead defense, and coupling trap—and it seems the defenses are just different ways of trying to manage that safety versus capability balance <ref:2605.14605#pg0>.
Priya: It's interesting how these templates show that almost every attempt to preserve utility has an inherent trade-off with the ability to stop harmful behavior from emerging.
Nadia: The paper then introduces a unified adaptive attack called SIDESTEPPER, which breaks all those defense methods by optimizing a full success criterion instead of just focusing on harmfulness <ref:2605.14605#pg0>.
Elias: That mixed objective, Latk(θ) = Lh(θ) + λLc(θ), which balances harmful behavior loss with benign capability loss, seems like the perfect counter because it steers the search toward parameters that are both harmful and useful <ref:2605.14605#pg0>.
Priya: From a data perspective, this attack isn't just looking for toxicity; it’s actively searching for a model state where harm is present but utility hasn't completely vanished, which is a much more realistic threat model.
Nadia: The results are telling because they show that this single adaptive objective breaks all the defense methods tested, suggesting the vulnerability isn't in any one defense design but in the common assumption about what an attacker cares about <ref:2605.14605#pg0>.
Elias: That really hammers home that we need to stop defending against simple harmful fine-tuning and start testing against this combined objective function.
Priya: If this is true, it means the models we release might pass safety checks designed for static attacks, but still be vulnerable to these more sophisticated objectives in practice.
Nadia: So, what does the paper suggest we actually do about improving these defenses moving forward? The authors point toward a necessary shift in how we evaluate robustness <ref:2605.14605#pg1>.
Title and authors: Elias: They emphasize that evaluation needs to include testing against adversaries who deliberately change the optimization objective to bypass existing defenses, which is what they call adaptive evaluation <ref:2605.14605#pg2>.
Priya: I’m hoping this leads to better measurement standards for safety; we need benchmarks that reflect this more complex, multi-objective optimization landscape rather than just single metrics like harmful loss alone.
Nadia: The main improvement suggested is that any future MFT defense should report robustness against an attacker minimizing the joint objective of harm plus benign capability <ref:2605.14605#pg0>.
Elias: That’s a very concrete prescription; we stop testing against just Lh and start testing against Lh plus a weighted component of Lc, which is the model's usefulness <ref:2605.14605#pg0>.
Priya: It sounds like the future involves developing new evaluation frameworks that specifically test how models maintain capability when subjected to these dual-objective optimization scenarios.
Nadia: It really does; the paper concludes that robustness against Lh-only fine-tuning is not evidence of robustness against malicious fine-tuning, and we need to adopt this joint objective as the minimum bar for security <ref:2605.14605#pg1>.
Elias: So, the implications are clear: defenses must be designed with the expectation that an adversary will optimize for both harmfulness and utility simultaneously during fine-tuning <ref:2605.14605#pg2>.
Priya: It suggests that we can’t just focus on blocking specific harmful outputs; we have to ensure the underlying mechanism stays functional across a wider, more complex optimization space.
Nadia: To wrap things up, the paper "A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training" shows that defenses fail because they are evaluated against fixed attacks when the real threat involves adaptive adversaries optimizing for both harm and utility <ref:2605.14605#pg0>.
Elias: It really underscores the need to move beyond simple harm detection and consider the full training objective, which is what this research highlights <ref:2605.14605#pg1>.
Priya: This means our work in measurement has to evolve to test models under these more realistic, multi-objective attack scenarios instead of just single-vector evaluations <ref:2605.14605#pg2>.
Nadia: It’s a call for better testing protocols so we don't give the impression that a model is safe just because it passed one type of fine-tuning test <ref:2605.14605#pg1>.
Elias: We certainly need to keep our eyes on how the cryptographic proofs and assumptions in these defenses hold up when they are subjected to this kind of adaptive pressure <ref:2605.14605#pg2>.
Priya: This paper gives us a clear direction for improving the security landscape by focusing on the joint optimization goal rather than isolated safety metrics <ref:2605.14605#pg0>.
The paper's summary: Nadia: So, to recap, this paper is showing that current defenses against malicious fine-tuning aren't actually robust because they are designed to stop fixed attacks rather than adaptive ones that change their strategy while training forty.
Elias: Exactly; it’s about how those defenses only work against one specific type of adversary and don't account for the fact that a smart attacker can learn and adjust their approach over time.
Priya: From what I see in the summary, this research points out that almost all existing defense methods share a fundamental weakness: they manage to hide harmful behavior without actually eliminating the behavior itself <ref:2605.14605#pg0>.
Nadia: That’s a crucial distinction because it means these techniques are just redirecting the search path, not stopping the underlying malicious process from succeeding.
Elias: And that leads us into how they structure their defenses, which we see broken down into strategies like anchoring or self-destruction based on how they shape the loss landscape <ref:2605.14605#pg0>.
Priya: I’m interested in the way they decompose everything into those four loss templates—robust-alignment basin, harmful-information removal, look-ahead defense, and coupling trap—to show how these different defenses are just different ways of balancing safety versus capability <ref:2605.14605#pg0>.
Nadia: It sounds like the paper is laying out the architecture of these existing defenses so we can see exactly where they are weakest when faced with a more complex adversary forty.
Elias: And then they introduce SIDESTEPPER, which uses a mixed objective function that combines harmful loss and benign capability loss to create an adaptive attack <ref:2605.14605#pg0>.
Priya: That combined objective, Latk(θ) = Lh(θ) + λLc(θ), is what really gets the point across because it assumes the attacker cares about both harm and usefulness simultaneously <ref:2605.14605#pg0>.
Nadia: So, the core finding is that this mixed objective attack breaks every single defense method tested, suggesting that we need to stop thinking about defending against harm in isolation forty.
Elias: That implies the vulnerability isn't in a single defense design but stems from a common assumption about how an attacker optimizes their goals during training <ref:2605.14605#pg0>.
Priya: And this has huge implications for privacy and measurement because it means our current safety evaluations are likely passing models that are still highly susceptible to these more realistic, multi-objective optimization scenarios <ref:2605.14605#pg2>.
Nadia: It really highlights a gap in how we measure robustness; we need to move beyond just checking for harmful loss recovery and start testing against this joint objective of harm and utility preservation forty.
Elias: So, the paper suggests that the minimum standard for MFT security should be defined by defending against an attacker minimizing both Lh plus a weighted component of Lc <ref:2605.14605#pg0>.
Priya: That gives us a clear target for future research: we need new evaluation frameworks that test how models maintain capability when subjected to these dual-objective optimization scenarios <ref:2605.14605#pg2>.
Nadia: It’s a call for a fundamental shift in how we test, moving away from static defense checks toward dynamic adversarial objectives forty.
Elias: We also see the discussion around escape signals and trajectory awareness, showing that defenses only constrain a local region of the loss surface, which an adaptive attack can easily exploit <ref:2605.14605#pg1>.
Priya: It makes me think about how this could impact real-world deployment; if these defenses are brittle, we need to be very cautious about deploying models that rely on them for safety <ref:2605.14605#pg1>.
The paper's improvements: Tom: So, we're shifting gears now to what the authors suggest we actually do about fixing these brittle defenses, and this paper lays out some specific improvements <ref:2605.14605#pg1>.
Nadia: It sounds like the main takeaway is that any future MFT defense needs to report robustness against an attacker minimizing both harmful loss and benign capability loss, not just harm alone forty.
Elias: That's a big change because it forces the security community to adopt this joint objective as the actual minimum requirement for security in this space <ref:2605.14605#pg1>.
Priya: I’m interested in those suggestions regarding schedule-aware fine-tuning, specifically how incorporating dynamic learning rate adjustments can counter attacks that use learning rate "shocks" to bypass defenses <ref:2605.14605#pg3>.
Nadia: That makes sense because the KICK-SETTLE attack showed that schedule manipulation is a real way to get around those local constraints we talked about before forty.
Elias: And they suggest that for more complex models, we should look into trajectory-aware regularization, using a look-ahead loss term to penalize trajectories that lead toward capability collapse <ref:2605.14605#pg3>.
Priya: That speaks directly to the coupling trap and look-ahead defense templates; it suggests we need mechanisms that monitor the entire fine-tuning trajectory, not just a single step <ref:2605.14605#pg3>.
Nadia: So, in simple terms, they’re telling us we can't just build defenses that work against one kind of attack; we need to design systems that are resilient across the entire training process forty.
Elias: The idea is to move from localized protection to a system that anticipates the attacker's full optimization strategy by incorporating capability preservation into every step of the defense mechanism <ref:2605.14605#pg0>.
Priya: This means our measurement tools need to evolve significantly, focusing on testing models under these multi-objective attack scenarios instead of just single metrics like harmful loss recovery <ref:2605.14605#pg2>.
Nadia: It’s a call for better testing protocols so we don't give the impression that a model is safe just because it passed one type of fine-tuning test forty.
Elias: We also see they emphasize the need for adaptive evaluation, meaning researchers have to test defenses against adversaries who actively change their objective during training <ref:2605.14605#pg2>.
Priya: If this is implemented properly, it could lead to much more trustworthy model releases because we’d be testing against what constitutes a compromised model in practice, not just theoretical boundaries <ref:2605.14605#pg1>.
Conclusion: Tom: So, to wrap things up, this paper shows that robustness against Lh-only fine-tuning isn't really protection against malicious fine-tuning because defenses are just tested against fixed attacks forty.
Nadia: Exactly; the main implication is that we need to fundamentally change our evaluation bar for MFT security by requiring models to survive attacks optimized for both harm and utility simultaneously <ref:2605.14605#pg1>.
Elias: I agree; it means the cryptographic assumptions underpinning many of these defenses are being tested against an objective that is much harder to achieve, which points to a weakness in how we calculate those proofs forty.
Priya: From a privacy and measurement standpoint, this suggests that our current safety evaluations are likely passing models that remain highly susceptible to these adaptive objectives in real-world deployment scenarios <ref:2605.14605#pg2>.
Nadia: It really highlights a gap in how we measure robustness; we need to move beyond just checking for harmful loss recovery and start testing against this joint objective of harm and utility preservation forty.
Elias: That's a very concrete prescription; the minimum bar for security should be defined by defending against an attacker minimizing Lh plus a weighted component of Lc <ref:2605.14605#pg0>.
Priya: It sounds like the future involves developing new evaluation frameworks that specifically test how models maintain capability when subjected to these dual-objective optimization scenarios instead of just single-vector evaluations <ref:2605.14605#pg2>.
Nadia: It’s a call for a fundamental shift in how we test, moving away from static defense checks toward dynamic adversarial objectives forty.
Elias: We certainly need to keep our eyes on how the cryptographic proofs and assumptions in these defenses hold up when they are subjected to this kind of adaptive pressure forty.
Priya: This paper gives us a clear direction for improving the security landscape by focusing on the joint optimization goal rather than isolated safety metrics <ref:2605.14605#pg0>.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel