Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection
cs.CR, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Project page: https://christophm.github.io/interpretable-ml-book
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks.
Terminology
Abstract
Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against such attacks, but their internal decision logic is largely opaque to both defenders and attackers. This paper presents an exploratory case study that applies explainable artificial intelligence (XAI) techniques to analyze how Prompt Guard 2 distinguishes malicious from benign prompts. We conduct four experiments to probe this question empirically. Guided by Vanilla Gradient and SHAP attributions, we find that Prompt Guard 2's decisions rely on the cumulative contribution of many tokens rather than a few dominant ones, yet saliency-guided synonym substitution and sentence-level paraphrasing can flip its predictions while altering only a moderate fraction of the text, in some cases yielding a successful jailbreak against the underlying LLM. A dataset-scale saliency analysis further shows that undetected injection prompts systematically lack the lexical markers the classifier relies on. We discuss the implications of these findings for the design and evaluation of classifier-based guardrails, and argue that explanation methods intended to support transparency can simultaneously lower the cost of constructing successful adversarial bypasses.
Sources
- Towards A Rigorous Science of Interpretable Machine Learning
- Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- Very Deep Convolutional Networks for Large-Scale Image Recognition
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs