DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
cs.CR
Submitted: 2026-07-22
Updated: 2026-09-28
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- DynaGuard: A Dynamic Guardian Model With User-Defined Policies
- DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
- NonTextual Target Attack
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
- GuardReasoner: Towards Reasoning-based LLM Safeguards
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Granite Guardian
- DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
- Dynamic Jailbreaking Attack
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- Qwen3Guard Technical Report
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs