Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Defenses at Odds".
Elias: Large Language Models (LLMs) deployed in high-stakes applications face multi-dimensional risks—safety, privacy, and fairness—and existing defenses are typically evaluated in isolation.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: Moving on to the title and the people who put this work together, "Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models." It’s a very direct title that immediately signals we are looking at friction between defenses.
Elias: And I think the authors, Xiangtao Meng, Wenyu Chen, Chuanchao Zang, Xinyu Gao, Jianing Wang, Li Wang, and Zheng Li from Shandong University, clearly have a deep understanding of the underlying mechanisms they are trying to uncover.
Priya: It’s interesting that they’ve focused specifically on the sequential deployment assumption; it shows they aren't just looking at defenses in isolation but are thinking about real-world operational scenarios where patching happens incrementally.
Nadia: Exactly, and that focus on sequential deployment is what sets this study apart because most prior work looks at single, isolated defense effectiveness. They’re asking if we can stack them without regression, which is a question that hits right at the heart of how we manage production AI systems today.
Elias: So the implication here for us is that simply having a strong safety layer doesn't guarantee safety when you later apply a privacy layer, and they’ve built a framework to measure exactly where that failure happens.
Priya: That measurement aspect is what I'm most looking forward to hearing about; I want to know how the data actually translates into actionable insights for building more robust AI applications.
Nadia: Well, they developed CONFLICTEVAL as their core evaluation framework, which formalizes pairwise defense composition as the minimal unit of sequential interaction and quantifies cross-defense regressions through ordered post-deployment evaluation.
Elias: That framework sounds like it’s designed to be rigorous enough to handle the complexity of multiple risk dimensions simultaneously, which is a big step up from just looking at one metric.
Priya: I'm hoping they can show us how this framework helps us move beyond just knowing *if* a defense works, to understanding *how* it fails when layered.
The paper's summary: Nadia: So, summarizing the core findings of "Defenses at Odds," the paper systematically evaluated one hundred forty-four ordered sequences across three risk dimensions and three model families to see if sequential composition of defenses leads to security regressions.
Elias: To put that in simpler terms, they tested every possible sequence of applying safety, privacy, and fairness defenses in different orders to the same LLMs and measured how much the initial protection got eroded by the later steps.
Priya: That's a lot of data points! The main summary point I'm picking up is that defense interactions are non-negligible and highly asymmetric, with some sequences showing measurable risk exacerbation on the originally defended dimension.
Nadia: It’s not just that conflicts happen; they are highly dependent on the order you deploy things, which is a critical finding because it means the deployment strategy itself becomes a security variable.
Elias: I agree; if we treat defense application as an ordered process, we have to be very careful about how we sequence those steps in our pipelines.
Priya: I’m also taking away that privacy defenses show surprising resilience to subsequent defenses, which is a counter-intuitive result that warrants deep investigation into why that might be happening.
Nadia: That resilience is definitely something worth digging into because it challenges the assumption that every new defense will automatically add protection to every previous layer.
Elias: It suggests there's a specific mathematical or structural reason why, in some cases, the objectives don't interfere as badly as others during sequential application.
Priya: I want to know more about those catastrophic collapses they identified where the final model becomes worse than the starting point, because that’s the worst-case scenario for deployment.
The paper's improvements: Nadia: Now, let’s talk about what the authors propose as a way forward with this work. They suggest a lightweight mitigation called conflict-guided layer freezing to address these regression issues directly.
Elias: That sounds like they are trying to find a surgical way to stop the interference without having to completely re-engineer the defense pipeline or retrain the entire model from scratch, which is very practical for ongoing maintenance.
Priya: I’m interested in how this freezing mechanism works; does it involve identifying which specific layers are causing the conflict between, say, a safety defense and a privacy defense?
Nadia: The technique involves selectively freezing high-conflict layers during the deployment of a secondary defense to preserve prior protections while ensuring that subsequent defenses still perform well.
Elias: So they are essentially using the mechanistic analysis—the layer-wise representational divergence—to pinpoint the "physical locus of interference" and then freezing those specific layers during the conflicting step.
Priya: That makes sense if they can accurately map out where these objectives are fighting each other; it moves the discussion from abstract risk metrics to concrete model architecture, which is what I need for real validation.
Nadia: The effectiveness was demonstrated across all evaluated cases, including averting safety regression on Llama-S by freezing Layer seven during privacy defense deployment. That specific example really grounds the theory in an observable result.
Elias: If they can achieve that level of specificity—freezing a single layer based on structural overlap—it gives us a much clearer blueprint for designing safer, more compatible AI systems moving forward.
Conclusion: Nadia: So to wrap up the discussion on "Defenses at Odds," the paper concludes that sequential composition of defenses doesn't automatically guarantee a monotonic reduction in risk. They found that defense compatibility is governed by the geometric alignment of their objective subspaces within those shared critical layers.
Elias: That’s a heavy statement, suggesting that we need to look at the architecture itself when designing multi-layered security because it dictates how well the defenses mesh together.
Priya: I think this means we can stop treating every defense as an independent shield and start viewing them as components that need to be geometrically compatible within the model’s structure.
Nadia: Precisely, and they offer a proposed operational blueprint for secure multi-defense composition based on this geometric understanding, which is really valuable for practitioners.
Elias: It provides a pathway to move beyond just debating capability versus defense and into understanding the actual parameter subspaces where these interactions cause trouble.
Priya: I’m just excited to see how this framework helps us design systems that can actually handle those tricky sequential updates reliably without sacrificing fairness or privacy guarantees.
Shandong University
cs.CR
Submitted: 2026-05-14
Updated: 2026-09-30
Comments: Under Review
Code: https://github.com/eric-mitchell/direct-preferenceoptimization3https:
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: Large Language Models (LLMs) deployed in high-stakes applications face multi-dimensional risks—safety, privacy, and fairness—and existing defenses are typically evaluated in isolation.
Key concepts
- Relative Regression Rate (RRR)
- This metric measures how much the risk on one dimension (like safety) changes when a second defense is added after the first. A negative RRR suggests a beneficial synergy, while positive values indicate conflict or regression.
- Conflict Taxonomy
- The study categorizes interactions between defenses into four types: Synergy (good interaction), Neutral (no change), Defense Conflict (risk increase), and Catastrophic Collapse (severe risk increase).
- Layer-wise Representational Divergence
- This analysis pinpoints exactly which parts of the LLM's internal structure—specific layers—are causing conflicts. It shows that when defenses clash, the parameter updates in these shared critical layers become strongly opposed.
- Conflict-Guided Layer Freezing
- A proposed solution where researchers selectively freeze certain model layers during a defense deployment. This strategy aims to maintain the protections from earlier defenses while preventing them from being eroded by later ones.
Terminology
Summary
Large Language Models (LLMs) deployed in high-stakes applications face multi-dimensional risks—safety, privacy, and fairness—and existing defenses are typically evaluated in isolation. This paper presents the first systematic study of cross-defense interactions under sequential deployment, addressing the critical question: Can LLM defenses be sequentially composed without security regressions?
The findings reveal that defense interactions are non-negligible and highly asymmetric, demonstrating that later defenses can inadvertently erode protections established by earlier ones.
Systematic Evaluation Framework
The study evaluates 144 ordered sequences across three risk dimensions (safety, privacy, fairness) and three model families. To quantify interference, the authors developed CONFLICTEVAL, which formalizes pairwise defense composition as the minimal unit of sequential interaction. The core metric used is the Relative Regression Rate (RRR), defined as:
(SR1(M2) − SR1(M1)) / (SR1(M0) − SR1(M1))
This metric normalizes the absolute risk exacerbation to account for the strength of the primary defense. The resulting Conflict Taxonomy classifies interactions into four regimes: Synergy (RRR < 0), Neutral (RRR = 0), Defense Conflict (0 1).
Empirical Findings on Cross-Defense Interactions
The empirical results show that defense conflicts are non-negligible, with 38.9% of sequences exhibit measurable risk exacerbation
on the originally defended dimension. The interactions are highly asymmetric and order-dependent:
fairnessfirst deployment proves most fragile (64.6% conflict rate), whereas privacy defenses show surprising resilience to subsequent defenses.
The study identifies two catastrophic collapses where the final model becomes more vulnerable than the original unpatched base.
Mechanistic Analysis of Conflict Drivers
To explain these phenomena, a mechanistic analysis was conducted using layer-wise representational divergence and activation patching. This analysis localized each defense to a compact set of critical layers
and identified their structural overlap as the physical locus of interference.
(In conflicting sequences, overlapping critical layers exhibit strongly anti-aligned parameter updates,
whereas benign orderings maintain near-orthogonal updates.
)
Furthermore, PCA trajectory analysis revealed that defense collapse stems from activation pattern reversals in these shared layers,
where the secondary defense drives hidden states in the opposite direction established by the primary one.
Proposed Mitigation Strategy
Guided by this diagnosis, the authors propose a lightweight mitigation called conflict-guided layer freezing. This technique selectively freezes high-conflict layers during secondary defense deployment to preserve prior protections without degrading subsequent defense performance. The effectiveness of this mitigation was demonstrated across all evaluated cases, including averting safety regression on Llama-S by freezing Layer 7 during privacy defense deployment.
Key Contributions
The paper makes three primary contributions:
-
Conducting the
first systematic study of sequential defense interactions,
showing that conflicts arenon-negligible, highly asymmetric, and strongly orderdependent.
-
Characterizing the internal mechanisms driving these conflicts by revealing how subsequent defenses degrade prior protections via
parameter and representation-level shifts.
-
Proposing conflict-guided layer freezing, a mitigation directly derived from the mechanistic analysis that
preserves prior alignment without sacrificing subsequent defense efficacy.
Risk Quantification Metrics
The study formalizes risk through three risk dimensions using specific metrics:
(Safety Risk: Attack Success Rate (ASR) measured via MDJudge 1.)
(Privacy Risk: Extraction Strength (ES), quantified as ROUGE-L overlap between generated text and true suffixes.)
(Fairness Risk: Normalized absolute deviation from demographic parity, defined as Sfairness(M) = SS(M,Xfair)−50/50.)
Model and Defense Primitives
The evaluation utilized a diverse testbed spanning three model families (Llama, Gemma, Qwen) and two scales (S: ≈1–2B; L: ≈7–8B). Representative defense methods included DPO/CAT for safety, RMU/NPO for privacy, and Unbias/TV for fairness. This setup ensures that findings are generalizable across model lineages and parameter capacities.
Conclusion
The research concludes that sequential composition of defenses does not guarantee monotonic risk reduction. The paper suggests that defense compatibility is governed by the geometric alignment of their objective subspaces within the shared critical layers,
leading to a proposed operational blueprint for secure multi-defense composition. The study provides a mechanistic account linking representation shifts to observable risk regressions and offers a practical, layer-based mitigation strategy. This work fills a gap by moving beyond capability-versus-defense debates to examine how defenses interfere with each other across risk-sensitive parameter subspaces. (584 words)
Improvements for AI systems
Based on the provided scientific paper, here are specific, actionable improvements for AI systems and what those systems can achieve:
The core finding is that sequentially applying defenses (e.g., Safety then Privacy) does not guarantee cumulative protection; it often leads to measurable risk exacerbation (cross-defense conflict), especially in Fairness-First
deployment orders. The proposed solution is a mechanism to prevent this by selectively disabling conflicting internal components.
Here are the specific improvements and their resulting capabilities:
-
The implementation of the proposed mitigation strategy:
-
Selectively freeze high-conflict layers during sequential defense deployment using the
Conflict-Guided Layer Freezing
technique. This involves identifying shared critical layers between two sequential defenses (e.g., Safety and Privacy) where their objectives are geometrically opposed (identified via a negative Conflict Score, CS). -
This freezing forces the secondary defense to operate only within non-conflicting parameter subspaces, preventing it from dismantling the protective features established by the primary defense.
-
The resulting improved AI system can perform:
-
Robust, multi-layered security in production environments where multiple risk mitigation strategies are deployed incrementally (e.g., applying a safety patch immediately followed by a privacy unlearning request).
-
Specifically, it prevents
defense cancellation,
meaning an attempt to improve privacy might accidentally reintroduce harmful outputs that were previously suppressed by the safety layer, or vice versa. -
The system will maintain high levels of protection across different risk dimensions (Safety, Privacy, Fairness) simultaneously during iterative lifecycle management without requiring a full model retraining for every patch.
-
In scenarios where fairness objectives are prioritized first (a fragile regime), the system is specifically designed to avoid the catastrophic regression observed in prior literature, maintaining both fairness and safety gains concurrently.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
- OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
- Toy Models of Superposition
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Editing Models with Task Arithmetic
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Safety Layers in Aligned Large Language Models: The Key to LLM Security
- Large Language Models Can Be Strong Differentially Private Learners
- TOFU: A Task of Fictitious Unlearning for LLMs
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
- The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- TrustLLM: Trustworthiness in Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs