A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
cs.CR, cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 8 pages (main), with appendix
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful
Terminology
Abstract
Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse game, yet most remain static: their safety behavior is fixed at deployment, so they cannot accumulate defensive experience or adapt to unseen strategies. We propose a self-evolving test-time defense built around a persistent, cross-interaction rule memory: when an attack succeeds, the framework abstracts that failure into a method-level rule capturing the structural attack wrapper rather than the harmful topic, and reuses it against future inputs. Because rules are method-level, one induced rule generalizes across an entire attack family, and the label space expands as novel wrappers appear. The mechanism operates entirely through external memory and prompting, with no parameter updates, and applies to both open-weight and black-box API models. We realize it as four cooperating modules, but the contribution is the memory-based adaptation mechanism, not the module decomposition. Across four black-box jailbreak families and multiple models, our method substantially reduces attack success rates while preserving benign utility, remains robust under an adaptive composite-wrapper attack, and does not increase over-refusal as the memory grows.
Sources
- Jailbreaking Black Box Large Language Models in Twenty Queries
- The Llama 3 Herd of Models
- MART: Improving LLM Safety with Multi-round Automatic Red-Teaming
- A Survey on LLM-as-a-Judge
- Gemini: A Family of Highly Capable Multimodal Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
- AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
- Qwen2.5 Technical Report
- CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
- GPT-4 Technical Report
- GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
- SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
- A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
- Qwen3 Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs