When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
Tong Zhang, Zexin Li, Simin Chen, Yun Peng
cs.CR, cs.LG
Submitted: 2026-07-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility.
Terminology
Abstract
Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost. Rather than treating defenses as a single class, we organize them by operational strategy and examine how different strategies correlate with different side-effect profiles. Across state-of-the-art defense methods, widely used benchmark datasets, and representative open-source LLMs, we find that defenses rarely improve downstream capability, but instead vary in how they trade safety gains against usability and efficiency. In particular, rule-based defenses best preserve task performance, highly conservative self-reflective defenses often increase over-refusal, and multi-round defenses incur the largest runtime overhead. These results provide both a benchmark for evaluating defense side effects and practical guidance for selecting defenses under deployment constraints.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- GPT-4 Technical Report
- Detecting Language Model Attacks with Perplexity
- Language Models are Few-Shot Learners
- Training Verifiers to Solve Math Word Problems
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- The Llama 3 Herd of Models
- TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Mistral 7B
- Counting function estimates for coherent frames and Riesz sequences
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- LLM4Decompile: Decompiling Binary Code with Large Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
- Qwen2 Technical Report
- Intention Analysis Makes LLMs A Good Jailbreak Defender
- Instruction-Following Evaluation for Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs