Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense
cs.CR, cs.CL
Submitted: 2026-06-04
Updated: 2026-09-05
Comments: EMNLP 2026 Main
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks.
Terminology
Abstract
Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evolving attacks, while adaptive memory-based guardrails tend to over-refuse benign queries that resemble stored attacks. We propose Membrane, a self-evolving guardrail built on Contrastive Safety Memory (CSM): each cell pairs the conditions for blocking a harmful query with those for permitting a superficially similar benign request. Without retraining, Membrane evolves CSM by distilling each harmful interaction and its benign counterpart into a contrastive cell indexed by the underlying attack strategy, so that one cell generalizes across topical variants of the same mechanism. At inference, retrieved cells serve as grounding context for precise safety decisions. Across model-level safety on HarmBench and agent-level safety on AgentHarm, Membrane achieves the highest F1 on all six modern jailbreak attacks. Notably, benign refusal on AgentHarm stays at 7-14%, well below the 28-85% range of prior guards. Memory cells also retain 87-88% F1 under cross-attack transfer and remain stable under memory poisoning.
Sources
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- Adversarial D'ej\`a Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Passage Re-ranking with BERT
- Olmo 3
- Qwen3 Technical Report
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs