How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions
Yukai Zhou, Feiyang Lu, Xiaokai Mao, Jinfei Liu, Wenjie Wang
cs.CR, cs.CL
Submitted: 2026-07-19
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
- Robust LLM safeguarding via refusal feature adversarial training
- Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
- AdvPrefix: An Objective for Nuanced LLM Jailbreaks
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs