PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift
cs.CR, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 21 pages, 3 figures, 18 tables
Journal ref: IEEE Access, vol. 14, pp. 134184-134205, 2026
DOI: 10.1109/ACCESS.2026.3728186
Code: https://github.com/ShirePyDev/Prompt-Injection-Detection-System
License: http://creativecommons.org/licenses/by/4.0/
The gist: Prompt-injection detectors are typically evaluated using aggregate F1 on in-distribution test data, which offers limited insight into behavior under distribution shift, particularly on the benign
Terminology
Abstract
Prompt-injection detectors are typically evaluated using aggregate F1 on in-distribution test data, which offers limited insight into behavior under distribution shift, particularly on the benign side of the decision boundary, where false positives impose direct operational cost yet are seldom measured. We present PIDS-Bench, a frozen multi-axis benchmark that jointly evaluates attack detection and benign false-positive behavior at fixed thresholds, spanning in-distribution inputs, hard-benign prompts that mimic injection structure without malicious intent, obfuscated attacks, and domain and structural distribution shifts. We evaluate seven detectors (learned baselines, external prompt-injection classifiers, and broad-safety comparators) alongside a rule-based lower-bound reference. Multi-axis evaluation exposes a failure mode that aggregate F1 conceals. A detector exceeding F1 = 0.98 on the held-out split still misclassifies roughly one-third of an externally-sourced benign subset drawn from public corpora and restricted to security-adjacent content. Across a full threshold sweep and five training seeds, no internal detector reaches an operating point satisfying F1 >= 0.95 and hard-benign FPR <= 0.10 together on this stress distribution. Decomposing by provenance, we find that hard-negative augmentation nearly eliminates over-defense on curated stress inputs but leaves it substantially intact on externally-sourced prompts, a pattern we term provenance-sensitive over-defense. The asymmetry holds across both fine-tuned architectures and does not diminish as the augmentation pool grows, with the externally-sourced FPR remaining far above the 0.10 target. Whether augmentation matched to the externally-sourced distribution would close this gap is untested; threshold calibration and curated-style augmentation alone do not.
Sources
- Ignore Previous Prompt: Attack Techniques For Language Models
- Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Jailbroken: How Does LLM Safety Training Fail?
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Decoupled Weight Decay Regularization
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs