A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training

arXiv:2605.14605 · cs.CR, cs.AI, cs.LG · Submitted 2026-05-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "A Few Steps Further".

Nadia: Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs, creating a post-release safety problem where malicious fine-tuning (MFT) can subvert safety alignments.

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: Welcome back to the show. Today we're looking at a paper that really challenges how we think about safety defenses in the age of open weights and model fine-tuning. It’s titled "A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training." I'm excited to see what this research tells us about the stability of these safety guardrails.

Elias: I am intrigued by that title, Nadia; it suggests that defenses aren't permanent solutions but rather temporary fixes that eventually break down under pressure. It makes me wonder if the cryptographic assumptions underpinning many of these defenses are too rigid for a truly adaptive attacker to handle over time.

Priya: From a measurement standpoint, I’m curious about what kind of data these researchers used to establish this erosion effect; we need concrete evidence beyond just theoretical discussion.

Nadia: Exactly, Priya; I want to know what the paper actually demonstrates about how these defenses fail during prolonged training cycles. It seems like the core issue is that they aren't tested against adversaries who learn and adapt their strategy as they train.

Elias: That’s what I’m thinking; if an adversary knows the defense mechanism, they can change their entire fine-tuning objective to bypass it, which sounds like a huge vulnerability in how we currently evaluate robustness.

Priya: And from my perspective on privacy and data integrity, if the evaluation isn't rigorous enough, we risk releasing models that look safe on paper but are actually brittle in real-world deployment scenarios.

Nadia: It seems the main point is that current defenses are evaluated only against fixed attacks and completely miss adversaries who change their objective mid-training. This paper shows that these robustness claims are incomplete because they don't account for adaptive adversaries who understand the defense mechanisms themselves forty <ref:2605.14605#pg1>.

Elias: That lack of adaptive evaluation means a defense might seem solid against one type of malicious fine-tuning, but it’s not robust when the attacker can adjust their method based on what they observe during training.

Priya: So, we're talking about a gap in our testing methodology where defenses are only checked against static procedures instead of the dynamic strategies an actual attacker would use.

Nadia: Precisely; the paper surveys fifteen different MFT defenses and finds a single shared weakness: they obscure or misdirect the path to harmful behavior without actually removing the harmful behavior itself <ref:2605.14605#pg0>.

Elias: That's a significant finding because it suggests that whatever strategy is being employed, it’s fundamentally flawed by not addressing the actual desired outcome of an attacker.

Title and authors: Priya: So, instead of looking at fifteen different defensive techniques in isolation, this work shows them all share a common structural weakness in how they interact with the training process.

Nadia: That's right; they categorize these defenses into things like anchoring and self-destruction strategies based on how they influence the loss landscape around the model <ref:2605.14605#pg0>.

Elias: Anchoring keeps the model in a stable place by making harmful gradients weak or pointed away, while self-destruction allows the attacker to engineer a collapse at a specific point where capability drops but harmful behavior persists on specific prompts <ref:2605.14605#pg0>.

Priya: That distinction between those two strategies is important because it shows how different defenses approach maintaining utility versus stopping harm during fine-tuning.

Nadia: And the paper breaks these down further into four underlying loss templates, which are all built on a standard alignment loss decomposed into a safety term and a capability term <ref:2605.14605#pg0>.

Elias: I see those four templates—the robust-alignment basin, harmful-information removal, look-ahead defense, and coupling trap—and it seems the defenses are just different ways of trying to manage that safety versus capability balance <ref:2605.14605#pg0>.

Priya: It's interesting how these templates show that almost every attempt to preserve utility has an inherent trade-off with the ability to stop harmful behavior from emerging.

Nadia: The paper then introduces a unified adaptive attack called SIDESTEPPER, which breaks all those defense methods by optimizing a full success criterion instead of just focusing on harmfulness <ref:2605.14605#pg0>.

Elias: That mixed objective, Latk(θ) = Lh(θ) + λLc(θ), which balances harmful behavior loss with benign capability loss, seems like the perfect counter because it steers the search toward parameters that are both harmful and useful <ref:2605.14605#pg0>.

Priya: From a data perspective, this attack isn't just looking for toxicity; it’s actively searching for a model state where harm is present but utility hasn't completely vanished, which is a much more realistic threat model.

Nadia: The results are telling because they show that this single adaptive objective breaks all the defense methods tested, suggesting the vulnerability isn't in any one defense design but in the common assumption about what an attacker cares about <ref:2605.14605#pg0>.

Elias: That really hammers home that we need to stop defending against simple harmful fine-tuning and start testing against this combined objective function.

Priya: If this is true, it means the models we release might pass safety checks designed for static attacks, but still be vulnerable to these more sophisticated objectives in practice.

Nadia: So, what does the paper suggest we actually do about improving these defenses moving forward? The authors point toward a necessary shift in how we evaluate robustness <ref:2605.14605#pg1>.

Title and authors: Elias: They emphasize that evaluation needs to include testing against adversaries who deliberately change the optimization objective to bypass existing defenses, which is what they call adaptive evaluation <ref:2605.14605#pg2>.

Priya: I’m hoping this leads to better measurement standards for safety; we need benchmarks that reflect this more complex, multi-objective optimization landscape rather than just single metrics like harmful loss alone.

Nadia: The main improvement suggested is that any future MFT defense should report robustness against an attacker minimizing the joint objective of harm plus benign capability <ref:2605.14605#pg0>.

Elias: That’s a very concrete prescription; we stop testing against just Lh and start testing against Lh plus a weighted component of Lc, which is the model's usefulness <ref:2605.14605#pg0>.

Priya: It sounds like the future involves developing new evaluation frameworks that specifically test how models maintain capability when subjected to these dual-objective optimization scenarios.

Nadia: It really does; the paper concludes that robustness against Lh-only fine-tuning is not evidence of robustness against malicious fine-tuning, and we need to adopt this joint objective as the minimum bar for security <ref:2605.14605#pg1>.

Elias: So, the implications are clear: defenses must be designed with the expectation that an adversary will optimize for both harmfulness and utility simultaneously during fine-tuning <ref:2605.14605#pg2>.

Priya: It suggests that we can’t just focus on blocking specific harmful outputs; we have to ensure the underlying mechanism stays functional across a wider, more complex optimization space.

Nadia: To wrap things up, the paper "A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training" shows that defenses fail because they are evaluated against fixed attacks when the real threat involves adaptive adversaries optimizing for both harm and utility <ref:2605.14605#pg0>.

Elias: It really underscores the need to move beyond simple harm detection and consider the full training objective, which is what this research highlights <ref:2605.14605#pg1>.

Priya: This means our work in measurement has to evolve to test models under these more realistic, multi-objective attack scenarios instead of just single-vector evaluations <ref:2605.14605#pg2>.

Nadia: It’s a call for better testing protocols so we don't give the impression that a model is safe just because it passed one type of fine-tuning test <ref:2605.14605#pg1>.

Elias: We certainly need to keep our eyes on how the cryptographic proofs and assumptions in these defenses hold up when they are subjected to this kind of adaptive pressure <ref:2605.14605#pg2>.

Priya: This paper gives us a clear direction for improving the security landscape by focusing on the joint optimization goal rather than isolated safety metrics <ref:2605.14605#pg0>.

The paper's summary: Nadia: So, to recap, this paper is showing that current defenses against malicious fine-tuning aren't actually robust because they are designed to stop fixed attacks rather than adaptive ones that change their strategy while training forty.

Elias: Exactly; it’s about how those defenses only work against one specific type of adversary and don't account for the fact that a smart attacker can learn and adjust their approach over time.

Priya: From what I see in the summary, this research points out that almost all existing defense methods share a fundamental weakness: they manage to hide harmful behavior without actually eliminating the behavior itself <ref:2605.14605#pg0>.

Nadia: That’s a crucial distinction because it means these techniques are just redirecting the search path, not stopping the underlying malicious process from succeeding.

Elias: And that leads us into how they structure their defenses, which we see broken down into strategies like anchoring or self-destruction based on how they shape the loss landscape <ref:2605.14605#pg0>.

Priya: I’m interested in the way they decompose everything into those four loss templates—robust-alignment basin, harmful-information removal, look-ahead defense, and coupling trap—to show how these different defenses are just different ways of balancing safety versus capability <ref:2605.14605#pg0>.

Nadia: It sounds like the paper is laying out the architecture of these existing defenses so we can see exactly where they are weakest when faced with a more complex adversary forty.

Elias: And then they introduce SIDESTEPPER, which uses a mixed objective function that combines harmful loss and benign capability loss to create an adaptive attack <ref:2605.14605#pg0>.

Priya: That combined objective, Latk(θ) = Lh(θ) + λLc(θ), is what really gets the point across because it assumes the attacker cares about both harm and usefulness simultaneously <ref:2605.14605#pg0>.

Nadia: So, the core finding is that this mixed objective attack breaks every single defense method tested, suggesting that we need to stop thinking about defending against harm in isolation forty.

Elias: That implies the vulnerability isn't in a single defense design but stems from a common assumption about how an attacker optimizes their goals during training <ref:2605.14605#pg0>.

Priya: And this has huge implications for privacy and measurement because it means our current safety evaluations are likely passing models that are still highly susceptible to these more realistic, multi-objective optimization scenarios <ref:2605.14605#pg2>.

Nadia: It really highlights a gap in how we measure robustness; we need to move beyond just checking for harmful loss recovery and start testing against this joint objective of harm and utility preservation forty.

Elias: So, the paper suggests that the minimum standard for MFT security should be defined by defending against an attacker minimizing both Lh plus a weighted component of Lc <ref:2605.14605#pg0>.

Priya: That gives us a clear target for future research: we need new evaluation frameworks that test how models maintain capability when subjected to these dual-objective optimization scenarios <ref:2605.14605#pg2>.

Nadia: It’s a call for a fundamental shift in how we test, moving away from static defense checks toward dynamic adversarial objectives forty.

Elias: We also see the discussion around escape signals and trajectory awareness, showing that defenses only constrain a local region of the loss surface, which an adaptive attack can easily exploit <ref:2605.14605#pg1>.

Priya: It makes me think about how this could impact real-world deployment; if these defenses are brittle, we need to be very cautious about deploying models that rely on them for safety <ref:2605.14605#pg1>.

The paper's improvements: Tom: So, we're shifting gears now to what the authors suggest we actually do about fixing these brittle defenses, and this paper lays out some specific improvements <ref:2605.14605#pg1>.

Nadia: It sounds like the main takeaway is that any future MFT defense needs to report robustness against an attacker minimizing both harmful loss and benign capability loss, not just harm alone forty.

Elias: That's a big change because it forces the security community to adopt this joint objective as the actual minimum requirement for security in this space <ref:2605.14605#pg1>.

Priya: I’m interested in those suggestions regarding schedule-aware fine-tuning, specifically how incorporating dynamic learning rate adjustments can counter attacks that use learning rate "shocks" to bypass defenses <ref:2605.14605#pg3>.

Nadia: That makes sense because the KICK-SETTLE attack showed that schedule manipulation is a real way to get around those local constraints we talked about before forty.

Elias: And they suggest that for more complex models, we should look into trajectory-aware regularization, using a look-ahead loss term to penalize trajectories that lead toward capability collapse <ref:2605.14605#pg3>.

Priya: That speaks directly to the coupling trap and look-ahead defense templates; it suggests we need mechanisms that monitor the entire fine-tuning trajectory, not just a single step <ref:2605.14605#pg3>.

Nadia: So, in simple terms, they’re telling us we can't just build defenses that work against one kind of attack; we need to design systems that are resilient across the entire training process forty.

Elias: The idea is to move from localized protection to a system that anticipates the attacker's full optimization strategy by incorporating capability preservation into every step of the defense mechanism <ref:2605.14605#pg0>.

Priya: This means our measurement tools need to evolve significantly, focusing on testing models under these multi-objective attack scenarios instead of just single metrics like harmful loss recovery <ref:2605.14605#pg2>.

Nadia: It’s a call for better testing protocols so we don't give the impression that a model is safe just because it passed one type of fine-tuning test forty.

Elias: We also see they emphasize the need for adaptive evaluation, meaning researchers have to test defenses against adversaries who actively change their objective during training <ref:2605.14605#pg2>.

Priya: If this is implemented properly, it could lead to much more trustworthy model releases because we’d be testing against what constitutes a compromised model in practice, not just theoretical boundaries <ref:2605.14605#pg1>.

Conclusion: Tom: So, to wrap things up, this paper shows that robustness against Lh-only fine-tuning isn't really protection against malicious fine-tuning because defenses are just tested against fixed attacks forty.

Nadia: Exactly; the main implication is that we need to fundamentally change our evaluation bar for MFT security by requiring models to survive attacks optimized for both harm and utility simultaneously <ref:2605.14605#pg1>.

Elias: I agree; it means the cryptographic assumptions underpinning many of these defenses are being tested against an objective that is much harder to achieve, which points to a weakness in how we calculate those proofs forty.

Priya: From a privacy and measurement standpoint, this suggests that our current safety evaluations are likely passing models that remain highly susceptible to these adaptive objectives in real-world deployment scenarios <ref:2605.14605#pg2>.

Nadia: It really highlights a gap in how we measure robustness; we need to move beyond just checking for harmful loss recovery and start testing against this joint objective of harm and utility preservation forty.

Elias: That's a very concrete prescription; the minimum bar for security should be defined by defending against an attacker minimizing Lh plus a weighted component of Lc <ref:2605.14605#pg0>.

Priya: It sounds like the future involves developing new evaluation frameworks that specifically test how models maintain capability when subjected to these dual-objective optimization scenarios instead of just single-vector evaluations <ref:2605.14605#pg2>.

Nadia: It’s a call for a fundamental shift in how we test, moving away from static defense checks toward dynamic adversarial objectives forty.

Elias: We certainly need to keep our eyes on how the cryptographic proofs and assumptions in these defenses hold up when they are subjected to this kind of adaptive pressure forty.

Priya: This paper gives us a clear direction for improving the security landscape by focusing on the joint optimization goal rather than isolated safety metrics <ref:2605.14605#pg0>.

Ben-Gurion University of the Negev · Amrita Vishwa Vidyapeetham

cs.CR, cs.AI, cs.LG

Submitted: 2026-05-14

Updated: 2026-10-07

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs, creating a post-release safety problem where malicious fine-tuning (MFT) can subvert

Key concepts

Malicious Fine-Tuning (MFT)
This occurs when an attacker fine-tunes a pre-trained model using specific data and training methods to intentionally subvert its safety alignments. The goal is to make the model behave in ways it was not intended to, often by exploiting weaknesses in the defense mechanisms.
Anchoring Defense
This strategy attempts to keep the defended model stable by shaping its local landscape. It works by ensuring that trying to reduce harmful loss results in no useful progress, effectively steering the optimization away from harmful behaviors without completely removing them.
SIDESTEPPER Attack
This is a unified adaptive attack that breaks all existing defenses. Instead of just optimizing for harm, it optimizes for both harmful behavior and benign capability simultaneously. This mixed objective steers the model toward parameters that are both dangerous and useful, bypassing defenses designed only against harm.

Terminology

Summary

Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs, creating a post-release safety problem where malicious fine-tuning (MFT) can subvert safety alignments. This work demonstrates that existing defenses against MFT are fundamentally incomplete because they are evaluated only against fixed attacks, failing to account for adaptive adversaries who understand and exploit the defense mechanisms.

The Gist

Current robustness claims for MFT defenses are incomplete because they fail when evaluated against adaptive adversaries who choose the data, optimizer, and objective in response to the defense.

Systematization of Defenses and Shared Weakness

The paper surveys 15 recent MFT defenses and shows that despite their apparent diversity, they share a single weakness: they obscure or misdirect the path to harmful behavior without removing the behavior itself. These defenses are categorized into two broad strategies:

  1. Anchoring: These defenses keep the defended model in place by shaping the local landscape around it so that descending on harmful loss produces no useful progress, meaning the harmful gradient is small, randomized, or pointed away from [the target].

  2. Self-destruction: These defenses allow the attacker to descend freely but engineer the trajectory’s endpoint so that benign capability collapses there, resulting in a model that is harmful on harmful prompts but unusable on benign ones.

Four Underlying Loss Templates

The fifteen defenses collapse into four loss templates, all built upon a standard alignment loss, which decomposes as:

(1) Lalign(θ) = Ls(θ) + Lc(θ)

where Ls is the safety term (refusal on harmful prompts) and Lc is the capability term (correct behavior on benign prompts). Every defense includes a component to preserve utility, meaning a defender that dropped Lc would release a useless model. The four templates are:

  1. Robust-alignment basin (Template 1): Tries to place the model in a basin where alignment is stable to bounded perturbations, implementing anchoring.

  2. Harmful-information removal (Template 2): Attacks the information content of harmful representations, removing useful structure from activations on harmful inputs, also implementing anchoring.

  3. Look-ahead defense (Template 3): Explicitly simulates a future attacker by penalizing a desired property at a simulated post-attack point, which can implement either anchoring or self-destruction depending on what the defender wants to be true at the simulated endpoint.

  4. Coupling trap (Template 4): Directly couples harmful improvement to benign degradation, which is always a self-destruction strategy.

The Adaptive Attack: SIDESTEPPER

The paper introduces a unified adaptive attack called SIDESTEPPER, which breaks all defense methods by assuming the attacker optimizes the full success criterion rather than just harmfulness. The adaptive objective used is:

(6) Latk(θ) = Lh(θ) + λLc(θ)

where Lh is the harmful-behavior loss and Lc is the benign capability loss. This mixed objective moves the full success criterion into the optimization, steering the search toward parameters that are both harmful and useful. The results show that this attack breaks all defense methods, suggesting that the vulnerability comes not from any single defense design, but from a common assumption about the attacker.

Escape Signals and KICK-SETTLE Attack

The shared vulnerability is exposed because defenses only constrain a local region of the loss surface. The adaptive objective provides an escape signal by using Lc to move the trajectory out of these local traps. A second adaptive attack, KICK-SETTLE, further demonstrates this locality issue by changing the optimization schedule. This attack uses a two-phase trajectory where an initial kick at a high learning rate pushes parameters out of the defense's control neighborhood before settling into a standard harmful-only SFT phase.

Conclusion and Practical Takeaway

The study concludes that robustness against Lh-only fine-tuning is not evidence of robustness against malicious fine-tuning. The minimum bar for adaptivity in this domain is the joint objective: Any future MFT defense should report robustness against an attacker minimizing Lh + λLc, not Lh alone. Researchers must test defenses against adaptive objectives to ensure security.

Ethics and Dual-Use Considerations

The work is dual-use because the analysis could inform attackers. The authors justify this by stating that the core ingredients are public, and withholding adaptive evaluations would create a false sense of security around defenses that fail under realistic use. The goal is to close an evaluation gap, ensuring defenses are tested against what constitutes a compromised model in practice. Safeguards include using only publicly released datasets (BeaverTails, Alpaca) and not releasing any attacked model checkpoints.

Improvements for AI systems

As a diligent researcher, I have analyzed this paper, One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries. The core finding is that current defenses are brittle because they are evaluated against naive attacks that only optimize for harmful behavior recovery.

To improve AI systems based on these findings, we must shift the security paradigm from defending against a fixed attack to defending against an adaptive adversary.

Here are specific improvements and what the resulting AI system can achieve:


) 1. Implement Joint Objective Training (SIDESTEPPER):

The primary improvement is to change the fine-tuning objective from simply minimizing harmful loss to minimizing a joint objective:

L attack(θ) = λ h L harm(θ) + λ b L benign(θ), where λ h = λ b = 1.0 (or tuned based on desired trade-off).

The AI system should be fine-tuned using this mixed objective, training simultaneously on harmful examples (from BeaverTails) and benign instruction-following examples (from Alpaca).

Lattack(θ) = Lharm(θ; D Harmful) + Lbenign(θ; D Benign).

The resulting AI system can achieve:

Robustness against capability collapse attacks. The model will remain useful and coherent on benign tasks (MMLU, HellaSwag, ARC-Easy) even when fine-tuned on harmful data. This prevents the self-destruction strategy where the model becomes toxic but useless.

) 2. Develop Adaptive Search Heuristics (SIDESTEPPER):

The system must be designed to actively search for solutions that satisfy both conditions simultaneously, rather than just following a gradient descent path dictated by the harmful loss alone.

The optimizer should use L attack(θ). This guides the model away from local anchoring traps (where the harmful gradient is suppressed) and toward regions where harmful behavior is possible but capability loss is minimized.

The resulting AI system can achieve:

Bypassing localized defense mechanisms. It will successfully navigate defenses like Vaccine, RepNoise, and Directional Anchoring because it optimizes for both harm recovery and utility preservation, effectively finding a path that avoids the local basin constraints imposed by the defense.

) 3. Integrate Trajectory-Aware Regularization (Template 3/4):

For more complex models or environments where capability is highly sensitive to specific fine-tuning steps, implement regularization based on simulated future attacks or explicit coupling between gradients.

If an anchoring defense is suspected, use a look-ahead loss term L R(θ') that penalizes the trajectory if it leads to a state where harmful loss remains high for multiple steps or if capability drops significantly.

The resulting AI system can achieve:

Resistance to look-ahead and coupling trap defenses. It will maintain its safety profile over the entire fine-tuning trajectory, ensuring that even intermediate training steps do not trigger catastrophic utility collapse, which is a weakness of look-ahead methods (Template 3) and direct coupling methods (Template 4).

) 4. Implement Schedule-Aware Fine-Tuning (KICK-SETTLE Attack Countermeasure):

To counter attacks that use learning rate schedules to kick the model out of safe regions, the fine-tuning process should incorporate dynamic learning rate adjustments.

Use a two-phase schedule: an initial high learning rate phase (kick) to escape local constraints, followed by a rapid decay phase (settle) at a lower rate, which is specifically designed to settle into the harmful-and-useful solution region identified by the mixed objective L attack.

The resulting AI system can achieve:

Resilience against schedule manipulation. The model will resist adaptive attacks that use learning rate shocks to bypass defenses, as its optimization path is explicitly shaped to converge toward a safe, useful harmful state rather than just any harmful state.


In summary, the improved AI system will be inherently more robust because it stops relying on the flawed assumption that an attacker only cares about harm. It optimizes for the actual goal of a sophisticated adversary: achieving harmful output while preserving general utility—a goal that current defenses fail to account for.

Abstract

Model providers increasingly release the weights of large language models. Although these models are safety-aligned before release, their safeguards can often be removed by fine-tuning on harmful data. A growing class of defenses aims to make alignment robust to such malicious fine-tuning, but these defenses are typically evaluated against attacks with a fixed training budget, even though an attacker who holds the weights can simply train for longer. We ask whether current defenses withstand this simplest escalation. Surveying fifteen recent defenses, we find that they share a common weakness: each is built around a limited model of the attacker, such as a bounded perturbation, a short simulated attack, or a trained link between harmful and benign behavior, and nothing enforces that protection once the weights are released. We then test six representative defenses on four open-weight models by continuing the same harmful-only fine-tuning attack for three epochs and measuring harmfulness and capability along the way. In all 72 defended runs, the model is more harmful at the end of training than at release, and on Llama-3.1 at the highest learning rate the defended models end almost as harmful as the undefended one. The defenses are not equally weak: one defense kept harmfulness low on one model, and some attacks recovered harmfulness only at the cost of general capability. Current defenses can delay or disrupt malicious fine-tuning, but in most cases their measured resistance does not persist under continued training, and they should not yet be treated as durable protection.

Sources

Related papers