Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense
cs.CR, cs.LG
Submitted: 2026-05-29
Updated: 2026-08-26
Comments: We propose SCOUT, a detector allocation framework that predicts each detector's accuracy and latency on a given input before running it, letting operators control the safety-utility trade-off with a single threshold and route to an LLM judge only when needed
Project page: https://rockyli11.github.io/SCOUT
License: http://creativecommons.org/licenses/by/4.0/
The gist: Prompt-injection detectors are heterogeneous: each is strong on a different slice of attacks, and none is always reliable.
Terminology
Abstract
Prompt-injection detectors are heterogeneous: each is strong on a different slice of attacks, and none is always reliable. Yet existing systems still treat detection as a fixed single-detector pipeline, committing every request to one detector's blind spots. We reframe defense as detector allocation: given a heterogeneous pool, decide per request which detectors to run and whether to escalate to an LLM judge. Our framework SCOUT (Scalable and Controllable Outcome-prediction for Uncertainty-aware Triage) makes this decision dynamic by predicting each detector's per-sample reliability and latency from how it behaved on similar past inputs, and exposes a single safety-utility threshold to the operator (where utility bundles benign-pass rate and wall-clock). To evaluate this setting, we build SCOUT-450, a benchmark that captures the structurally complex, agent-facing injections that older prompt-injection sets under-represent. On SCOUT-450, a safety-oriented operating point reduces attack-success rate by 46% and total wall-clock by 40% relative to an always-on GPT-4o judge, at a 5.1-point benign-utility drop. SCOUT also transfers to three external benchmarks (BIPIA, IPI, and IHEval), improving the safety-utility frontier.
Sources
- On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
- Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Trust The Typical
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
- Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
- WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
- The Llama 3 Herd of Models
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
- RouteLLM: Learning to Route LLMs with Preference Data
- HybridFlow: A Flexible and Efficient RLHF Framework
- GPT-4o System Card
- gpt-oss-120b & gpt-oss-20b Model Card
- OpenAI GPT-5 System Card
- Ignore Previous Prompt: Attack Techniques For Language Models
- Qwen3 Technical Report
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs