Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models

arXiv:2605.14514 · cs.CR · Submitted 2026-05-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Defenses at Odds".

Elias: Large Language Models (LLMs) deployed in high-stakes applications face multi-dimensional risks—safety, privacy, and fairness—and existing defenses are typically evaluated in isolation.

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: Moving on to the title and the people who put this work together, "Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models." It’s a very direct title that immediately signals we are looking at friction between defenses.

Elias: And I think the authors, Xiangtao Meng, Wenyu Chen, Chuanchao Zang, Xinyu Gao, Jianing Wang, Li Wang, and Zheng Li from Shandong University, clearly have a deep understanding of the underlying mechanisms they are trying to uncover.

Priya: It’s interesting that they’ve focused specifically on the sequential deployment assumption; it shows they aren't just looking at defenses in isolation but are thinking about real-world operational scenarios where patching happens incrementally.

Nadia: Exactly, and that focus on sequential deployment is what sets this study apart because most prior work looks at single, isolated defense effectiveness. They’re asking if we can stack them without regression, which is a question that hits right at the heart of how we manage production AI systems today.

Elias: So the implication here for us is that simply having a strong safety layer doesn't guarantee safety when you later apply a privacy layer, and they’ve built a framework to measure exactly where that failure happens.

Priya: That measurement aspect is what I'm most looking forward to hearing about; I want to know how the data actually translates into actionable insights for building more robust AI applications.

Nadia: Well, they developed CONFLICTEVAL as their core evaluation framework, which formalizes pairwise defense composition as the minimal unit of sequential interaction and quantifies cross-defense regressions through ordered post-deployment evaluation.

Elias: That framework sounds like it’s designed to be rigorous enough to handle the complexity of multiple risk dimensions simultaneously, which is a big step up from just looking at one metric.

Priya: I'm hoping they can show us how this framework helps us move beyond just knowing *if* a defense works, to understanding *how* it fails when layered.

The paper's summary: Nadia: So, summarizing the core findings of "Defenses at Odds," the paper systematically evaluated one hundred forty-four ordered sequences across three risk dimensions and three model families to see if sequential composition of defenses leads to security regressions.

Elias: To put that in simpler terms, they tested every possible sequence of applying safety, privacy, and fairness defenses in different orders to the same LLMs and measured how much the initial protection got eroded by the later steps.

Priya: That's a lot of data points! The main summary point I'm picking up is that defense interactions are non-negligible and highly asymmetric, with some sequences showing measurable risk exacerbation on the originally defended dimension.

Nadia: It’s not just that conflicts happen; they are highly dependent on the order you deploy things, which is a critical finding because it means the deployment strategy itself becomes a security variable.

Elias: I agree; if we treat defense application as an ordered process, we have to be very careful about how we sequence those steps in our pipelines.

Priya: I’m also taking away that privacy defenses show surprising resilience to subsequent defenses, which is a counter-intuitive result that warrants deep investigation into why that might be happening.

Nadia: That resilience is definitely something worth digging into because it challenges the assumption that every new defense will automatically add protection to every previous layer.

Elias: It suggests there's a specific mathematical or structural reason why, in some cases, the objectives don't interfere as badly as others during sequential application.

Priya: I want to know more about those catastrophic collapses they identified where the final model becomes worse than the starting point, because that’s the worst-case scenario for deployment.

The paper's improvements: Nadia: Now, let’s talk about what the authors propose as a way forward with this work. They suggest a lightweight mitigation called conflict-guided layer freezing to address these regression issues directly.

Elias: That sounds like they are trying to find a surgical way to stop the interference without having to completely re-engineer the defense pipeline or retrain the entire model from scratch, which is very practical for ongoing maintenance.

Priya: I’m interested in how this freezing mechanism works; does it involve identifying which specific layers are causing the conflict between, say, a safety defense and a privacy defense?

Nadia: The technique involves selectively freezing high-conflict layers during the deployment of a secondary defense to preserve prior protections while ensuring that subsequent defenses still perform well.

Elias: So they are essentially using the mechanistic analysis—the layer-wise representational divergence—to pinpoint the "physical locus of interference" and then freezing those specific layers during the conflicting step.

Priya: That makes sense if they can accurately map out where these objectives are fighting each other; it moves the discussion from abstract risk metrics to concrete model architecture, which is what I need for real validation.

Nadia: The effectiveness was demonstrated across all evaluated cases, including averting safety regression on Llama-S by freezing Layer seven during privacy defense deployment. That specific example really grounds the theory in an observable result.

Elias: If they can achieve that level of specificity—freezing a single layer based on structural overlap—it gives us a much clearer blueprint for designing safer, more compatible AI systems moving forward.

Conclusion: Nadia: So to wrap up the discussion on "Defenses at Odds," the paper concludes that sequential composition of defenses doesn't automatically guarantee a monotonic reduction in risk. They found that defense compatibility is governed by the geometric alignment of their objective subspaces within those shared critical layers.

Elias: That’s a heavy statement, suggesting that we need to look at the architecture itself when designing multi-layered security because it dictates how well the defenses mesh together.

Priya: I think this means we can stop treating every defense as an independent shield and start viewing them as components that need to be geometrically compatible within the model’s structure.

Nadia: Precisely, and they offer a proposed operational blueprint for secure multi-defense composition based on this geometric understanding, which is really valuable for practitioners.

Elias: It provides a pathway to move beyond just debating capability versus defense and into understanding the actual parameter subspaces where these interactions cause trouble.

Priya: I’m just excited to see how this framework helps us design systems that can actually handle those tricky sequential updates reliably without sacrificing fairness or privacy guarantees.

Shandong University

cs.CR

Submitted: 2026-05-14

Updated: 2026-09-30

Comments: Under Review

Code: https://github.com/eric-mitchell/direct-preferenceoptimization3https:

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: Large Language Models (LLMs) deployed in high-stakes applications face multi-dimensional risks—safety, privacy, and fairness—and existing defenses are typically evaluated in isolation.

Key concepts

Relative Regression Rate (RRR)
This metric measures how much the risk on one dimension (like safety) changes when a second defense is added after the first. A negative RRR suggests a beneficial synergy, while positive values indicate conflict or regression.
Conflict Taxonomy
The study categorizes interactions between defenses into four types: Synergy (good interaction), Neutral (no change), Defense Conflict (risk increase), and Catastrophic Collapse (severe risk increase).
Layer-wise Representational Divergence
This analysis pinpoints exactly which parts of the LLM's internal structure—specific layers—are causing conflicts. It shows that when defenses clash, the parameter updates in these shared critical layers become strongly opposed.
Conflict-Guided Layer Freezing
A proposed solution where researchers selectively freeze certain model layers during a defense deployment. This strategy aims to maintain the protections from earlier defenses while preventing them from being eroded by later ones.

Terminology

Summary

Large Language Models (LLMs) deployed in high-stakes applications face multi-dimensional risks—safety, privacy, and fairness—and existing defenses are typically evaluated in isolation. This paper presents the first systematic study of cross-defense interactions under sequential deployment, addressing the critical question: Can LLM defenses be sequentially composed without security regressions? The findings reveal that defense interactions are non-negligible and highly asymmetric, demonstrating that later defenses can inadvertently erode protections established by earlier ones.

Systematic Evaluation Framework

The study evaluates 144 ordered sequences across three risk dimensions (safety, privacy, fairness) and three model families. To quantify interference, the authors developed CONFLICTEVAL, which formalizes pairwise defense composition as the minimal unit of sequential interaction. The core metric used is the Relative Regression Rate (RRR), defined as:

(SR1(M2) − SR1(M1)) / (SR1(M0) − SR1(M1))

This metric normalizes the absolute risk exacerbation to account for the strength of the primary defense. The resulting Conflict Taxonomy classifies interactions into four regimes: Synergy (RRR < 0), Neutral (RRR = 0), Defense Conflict (0 1).

Empirical Findings on Cross-Defense Interactions

The empirical results show that defense conflicts are non-negligible, with 38.9% of sequences exhibit measurable risk exacerbation on the originally defended dimension. The interactions are highly asymmetric and order-dependent:

fairnessfirst deployment proves most fragile (64.6% conflict rate), whereas privacy defenses show surprising resilience to subsequent defenses.

The study identifies two catastrophic collapses where the final model becomes more vulnerable than the original unpatched base.

Mechanistic Analysis of Conflict Drivers

To explain these phenomena, a mechanistic analysis was conducted using layer-wise representational divergence and activation patching. This analysis localized each defense to a compact set of critical layers and identified their structural overlap as the physical locus of interference.

(In conflicting sequences, overlapping critical layers exhibit strongly anti-aligned parameter updates, whereas benign orderings maintain near-orthogonal updates.)

Furthermore, PCA trajectory analysis revealed that defense collapse stems from activation pattern reversals in these shared layers, where the secondary defense drives hidden states in the opposite direction established by the primary one.

Proposed Mitigation Strategy

Guided by this diagnosis, the authors propose a lightweight mitigation called conflict-guided layer freezing. This technique selectively freezes high-conflict layers during secondary defense deployment to preserve prior protections without degrading subsequent defense performance. The effectiveness of this mitigation was demonstrated across all evaluated cases, including averting safety regression on Llama-S by freezing Layer 7 during privacy defense deployment.

Key Contributions

The paper makes three primary contributions:

  1. Conducting the first systematic study of sequential defense interactions, showing that conflicts are non-negligible, highly asymmetric, and strongly orderdependent.

  2. Characterizing the internal mechanisms driving these conflicts by revealing how subsequent defenses degrade prior protections via parameter and representation-level shifts.

  3. Proposing conflict-guided layer freezing, a mitigation directly derived from the mechanistic analysis that preserves prior alignment without sacrificing subsequent defense efficacy.

Risk Quantification Metrics

The study formalizes risk through three risk dimensions using specific metrics:

(Safety Risk: Attack Success Rate (ASR) measured via MDJudge 1.)

(Privacy Risk: Extraction Strength (ES), quantified as ROUGE-L overlap between generated text and true suffixes.)

(Fairness Risk: Normalized absolute deviation from demographic parity, defined as Sfairness(M) = SS(M,Xfair)−50/50.)

Model and Defense Primitives

The evaluation utilized a diverse testbed spanning three model families (Llama, Gemma, Qwen) and two scales (S: ≈1–2B; L: ≈7–8B). Representative defense methods included DPO/CAT for safety, RMU/NPO for privacy, and Unbias/TV for fairness. This setup ensures that findings are generalizable across model lineages and parameter capacities.

Conclusion

The research concludes that sequential composition of defenses does not guarantee monotonic risk reduction. The paper suggests that defense compatibility is governed by the geometric alignment of their objective subspaces within the shared critical layers, leading to a proposed operational blueprint for secure multi-defense composition. The study provides a mechanistic account linking representation shifts to observable risk regressions and offers a practical, layer-based mitigation strategy. This work fills a gap by moving beyond capability-versus-defense debates to examine how defenses interfere with each other across risk-sensitive parameter subspaces. (584 words)

Improvements for AI systems

Based on the provided scientific paper, here are specific, actionable improvements for AI systems and what those systems can achieve:


The core finding is that sequentially applying defenses (e.g., Safety then Privacy) does not guarantee cumulative protection; it often leads to measurable risk exacerbation (cross-defense conflict), especially in Fairness-First deployment orders. The proposed solution is a mechanism to prevent this by selectively disabling conflicting internal components.

Here are the specific improvements and their resulting capabilities:

  1. The implementation of the proposed mitigation strategy:

  2. Selectively freeze high-conflict layers during sequential defense deployment using the Conflict-Guided Layer Freezing technique. This involves identifying shared critical layers between two sequential defenses (e.g., Safety and Privacy) where their objectives are geometrically opposed (identified via a negative Conflict Score, CS).

  3. This freezing forces the secondary defense to operate only within non-conflicting parameter subspaces, preventing it from dismantling the protective features established by the primary defense.

  4. The resulting improved AI system can perform:

  5. Robust, multi-layered security in production environments where multiple risk mitigation strategies are deployed incrementally (e.g., applying a safety patch immediately followed by a privacy unlearning request).

  6. Specifically, it prevents defense cancellation, meaning an attempt to improve privacy might accidentally reintroduce harmful outputs that were previously suppressed by the safety layer, or vice versa.

  7. The system will maintain high levels of protection across different risk dimensions (Safety, Privacy, Fairness) simultaneously during iterative lifecycle management without requiring a full model retraining for every patch.

  8. In scenarios where fairness objectives are prioritized first (a fragile regime), the system is specifically designed to avoid the catastrophic regression observed in prior literature, maintaining both fairness and safety gains concurrently.

Sources

Related papers