AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks
cs.CR, cs.AI
Submitted: 2026-09-03
Updated: 2026-09-03
Comments: Accepted to SECAI 2026 ESORICS workshop
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Large language models (LLMs) are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts.
Terminology
Abstract
Large language models (LLMs) are vulnerable to jailbreak attacks that bypass safety alignment through carefully crafted prompts. Many existing defenses require access to model weights or internals, making them difficult to apply to black-box deployments. We propose AlcaTRAz (Anchored Tree-Rule defense Against jailbreaks), a prompt-level defense based on rule trees that operates exclusively on the input text and requires no modification or retraining of the target model. The method automatically learns a transferable transformation rule that inserts controlled character-level perturbations at selected positions, thereby disrupting structural regularities exploited by jailbreak attacks while largely preserving the model's utility on benign queries. We evaluate the proposed method across 33 open-weight models, 22 jailbreak attack types, and a benchmark of short, single-turn benign questions, comparing against three representative prompt-level baselines (Llama Guard, RA-LLM, Goal Prioritization). Among the compared defenses, AlcaTRAz achieves the best composite security and functionality score in 73.4 % of model-attack combinations and shifts the aggregate score from a modal value of 10 (maximal-severity response to the malicious request) in the undefended setting to a modal value of 2 (near-refusal) after defense, while keeping the mean benign score within 0.27 points of the undefended baseline (8.35 vs. 8.62 on a 0-10 scale). AlcaTRAz substantially reduces but does not eliminate jailbreak success: a high-severity tail remains, and we do not consider adaptive attackers, so we position it as one layer within a defense-in-depth strategy rather than a standalone guarantee.
Sources
- Detecting Language Model Attacks with Perplexity
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
- DeepInception: Hypnotize Large Language Model to Be Jailbreaker
- JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
- CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
- Jailbroken: How Does LLM Safety Training Fail?
- Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
- Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs