BRANCH: Bypassing Multi-Scanner AI Guardrails
summary
The gist
The gist: BRANCH, a bypassing methodology designed for multi-scanner guardrail systems, achieves 100% attack success rate across 6 guardrail systems in 120 scenarios with 72% fewer queries and 4.5x
In short
BRANCH is a novel method using a branching tree search to chain multiple evasion techniques simultaneously against multi-scanner AI guardrail systems. It successfully achieved a 100% attack success rate across six guardrail systems in 120 scenarios, significantly reducing query count and wall-clock time compared to existing methods while maintaining semantic meaning.
Key concepts
- Guardrail Systems
- These are sophisticated defenses that use an ensemble (combination) of multiple scanners to detect malicious content like prompt injection. Attackers must bypass every scanner in the system at once to succeed.
- Scanner Interference
- This occurs when combining the confidence scores from multiple scanners into one single score. This makes it harder for attackers because a specific behavior might look identical to another scanner, limiting how effective simple evasion techniques are against multi-scanner defenses.
- BRANCH Methodology
- This is a search strategy that uses a branching tree to chain different perturbation techniques together. It intelligently selects the next technique based on which one offers the lowest combined 'blocked confidence' across all scanners, guiding the attack toward simultaneous bypass.
Terminology used across episodes
This episode discusses
- BRANCH: Bypassing Multi-Scanner AI Guardrails · Paper Radio
- Generating Natural Language Adversarial Examples
- Current state of LLM Risks and AI Guardrails
- LlamaFirewall: An open source guardrail system for building secure AI agents
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning Attacks
- Safeguarding Large Language Models: A Survey
- The Llama 3 Herd of Models · Paper Radio
- Black-box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers
- Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
- SGuard-v1: Safety Guardrail for Large Language Models
- BERT-ATTACK: Adversarial Attack Against BERT Using BERT
- A Holistic Approach to Undesired Content Detection in the Real World
- Combating Adversarial Misspellings with Robust Word Recognition
- NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails
- An Early Categorization of Prompt Injection Attacks on Large Language Models
- Natural Language Adversarial Defense through Synonym Encoding
- A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- Prompt Overflow: What the Guardrail Inspects Is Not What the Model Infers
The paper
BRANCH: Bypassing Multi-Scanner AI Guardrails · Read on arXiv
William Hackett, Peter Garraghan
Mindgard · Lancaster University
AI systems increasingly rely on Large Language Models (LLMs) as core reasoning engines, making them targets for prompt injection and jailbreaks. Guardrails monitor and validate model inputs and outputs, yet their isolated, task-focused detection leaves gaps in their classification making them susceptible to bypasses. In response, guardrail systems formed by multiple scanners have emerged that collaboratively detect different types of malicious instructions, whereby shared latent representations across classification boundaries render established bypassing techniques ineffective. We propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems. Our method leverages a branching tree search approach that dynamically applies adversarial perturbation against individual scanners, with subsequent perturbation optimization and technique selection based on overall improvement across all guardrail system scanners, effectively decoupling bypass evaluation from attack signal optimization. Our findings demonstrate that BRANCH achieves 100% attack success rate across 6 guardrail systems in 120 scenarios with 72% fewer queries and 4.5x reduced wallclock time compared to established techniques, while preserving semantic meaning within the bypass. We also show how bypasses generated by BRANCH transfer to 29 unseen guardrails, including 8 commercial black-box guardrails, improving attack success in some cases up to 100% with no additional optimization.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "BRANCH: Bypassing Multi-Scanner AI Guardrails".
Nadia: The gist: BRANCH, a bypassing methodology designed for multi-scanner guardrail systems,
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: We’re diving into "BRANCH: Bypassing Multi-Scanner AI Guardrails," focusing on who wrote this and what the title actually suggests about the research direction. The authors are William Hackett and Peter Garraghan, working out of Mindgard at Lancaster University.
Elias: It immediately tells you that they are focused on moving beyond just single points of failure in AI security. The title itself highlights that it’s not about one scanner anymore, but about dealing with multiple scanners operating together as a guardrail system.
Nadia: Right, so the focus is clearly on multi-scanner environments. It’s not just looking at a single detection model anymore; it's about the entire ensemble of defenses working in concert against an attacker.
Priya: From my angle, when you see that title, I wonder if this suggests they are tackling a real weakness in how these guardrails are structured themselves, or if it’s just an evasion technique that happens to work better against ensembles.
Elias: It seems they're addressing the latter by proposing a methodology called BRANCH specifically designed for this multi-scanner challenge. It’s not just testing existing methods; it’s proposing a new framework for how to approach the bypass problem itself.
Nadia: Precisely, and what they do is propose this branching tree search structure. Instead of trying one big trick, you have multiple paths to explore until you find one that works against everything at once.
Priya: So, if we think about a listener driving and they hear this paper's title, what’s the simple translation? It means the AI safety net has multiple alarms ringing at once, and this method helps the attacker figure out how to silence all of them at the same time.
Elias: That’s a fair way to put it. They are looking for a way to navigate that chorus of detection systems without tripping any of their individual alerts in a destructive way.
The paper's summary: Nadia: Now let’s look at what the authors actually summarized in this paper about the BRANCH methodology itself. The main point is that they investigate how well single-objective evasion techniques fail when you try to bypass systems with multiple scanners.
Elias: They found that because these scanners share some underlying patterns or latent representations, trying to beat one scanner often gives you an advantage against another one in the system. This overlap is what makes established methods struggle in a multi-objective bypass scenario.
Priya: So, when they look at the numbers they present about scanner activations, what’s the most important thing for us to grasp about how many scanners are actually being triggered during these attacks?
Nadia: They show that across different scenarios, you can see a huge range in how many scanners are activated. For instance, in one set of tests, the number of scanners activated goes from zero up to over one thousand two hundred <ref:2610.10742#pg3>.
Elias: That variability shows how much the specific malicious instruction matters—some prompts barely activate anything, while others push all those detection layers simultaneously. It highlights that there isn't just one type of attack that works across all guardrails.
Priya: And what about the effectiveness of these different evasion techniques they mentioned, like genetic algorithms or BERT-based masking? Does the paper suggest any specific technique is inherently better for this multi-scanner problem?
Nadia: They investigate a few different ways to generate perturbations, but the key contribution is their selection strategy. BRANCH selects techniques based on how well they improve performance across *all* scanners combined, not just one.
Elias: So it’s not about finding the single best character substitution or word swap; it’s about managing a whole tree of possibilities to find a path that works across the entire detection setup.
The paper's improvements: Nadia: The authors propose several key improvements in this work, and one of the biggest is this novel multi-objective guardrail bypassing methodology they call BRANCH. It moves away from single-objective evasion by dynamically applying adversarial perturbations against individual scanners while constantly optimizing based on the total improvement across all scanners.
Priya: That sounds like a lot of coordination for the AI to manage, I guess? How does that dynamic application actually translate into a successful bypass in practice?
Elias: The paper demonstrates that this approach allows them to achieve a one hundred percent bypass success rate across six target guardrail systems <ref:2610.10742#pg2>. They’ve shown this capability without needing any additional manual optimization on their end for each new system.
Nadia: And the transferability is a huge part of it—they found that the bypasses generated by BRANCH successfully transfer to a large set of guardrails that were completely unseen when they first generated those prompts.
Priya: So, if I’m listening to this, what’s my takeaway about these improvements? It seems like the strength isn't in one clever trick, but in this systematic search process that keeps looking for a path through the entire defense structure.
Elias: Exactly. They found that this method is applicable to any malicious instruction they test, which means it’s not just a solution for one specific type of injection; it’s more general.
Conclusion: Nadia: So, to wrap up what we discussed about the "BRANCH: Bypassing Multi-Scanner AI Guardrails" paper, the main implication is that existing single-objective evasion techniques have a hard limit when facing multi-scanner guardrail systems. BRANCH provides a systematic way to overcome those limitations by searching for paths that satisfy all scanners simultaneously.
Elias: They achieved one hundred percent success across six systems with seventy-two percent fewer queries and a four point five times faster wall-clock time than established methods <ref:2610.10742#pg2>. That efficiency, combined with the transferability of the bypasses, is what they’re highlighting as important.
Priya: I think what stands out most for me is that it’s not just about attacking one system; it's about finding a generalized weakness in how these multi-layered guardrail systems are built. It suggests that defenses need to be designed with this kind of cross-system adversarial thinking in mind.
Nadia: That seems right. We get the idea that for AI security, we need to stop thinking about one scanner at a time and start thinking about the whole interconnected system when designing our defenses and our attacks.
Elias: It’s a solid piece of work showing how systematic search can actually be more effective than brute force in these complex adversarial scenarios. We should keep an eye on how this affects future guardrail design.
Priya: Well, that’s the summary for today on BRANCH: Bypassing Multi-Scanner AI Guardrails.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits