PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents
cs.CR, cs.CL
Submitted: 2026-06-16
Updated: 2026-09-12
Comments: 10 pages, 2 figures, 2 tables. Camera-ready version. Accepted to the 1st Workshop on Grounding Language Models: Learning Faithfully and Efficiently (GroundLM 2026) at EMNLP 2026
Code: https://github.com/aaditya79/parse-defense
License: http://creativecommons.org/licenses/by/4.0/
The gist: Prompt injection defenses evaluated on synthetic benchmarks do not generalize to real enterprise documents, which are longer, denser, and interleave legitimate authority language with factual content.
Terminology
Abstract
Prompt injection defenses evaluated on synthetic benchmarks do not generalize to real enterprise documents, which are longer, denser, and interleave legitimate authority language with factual content. We demonstrate this gap with a benchmark of 122 tasks across five professional domains (financial, legal, medical, scientific, DevOps) built on real retrieved documents -- actual SEC filings, Federal Register rules, PubMed abstracts, arXiv papers, and GitHub postmortems -- paired with LLM-generated tasks and camouflaged payloads that were not human-validated. Paraphrasing, the strongest defense on synthetic benchmarks, shows no statistically significant attack success rate reduction on real documents (p=0.500) while degrading utility from 91.8% to 82.8%. We introduce PARSE (Provenance-Aware Retrieval Sanitization), a domain-aware, fact-preserving sanitization pipeline that classifies each sentence by injection likelihood, extracts structured facts before rewriting, and verifies fact preservation via a consistency-checking loop. A directiveness gate routes 59% of real enterprise documents to a lightweight path, concentrating computational cost on high-risk documents. PARSE achieves 15.6% attack success rate -- a 39% reduction versus the 25.4% baseline -- at 86.9% utility, the largest reduction of any condition evaluated (h=-0.245), nominally significant (p=0.014) but not surviving correction for multiple comparisons. Practitioners should evaluate defenses on domain-matched real documents, not synthetic proxies.
Sources
- CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
- Evaluating Prompting-Based Defenses Against Domain-Camouflaged Injection Attacks
- Ignore Previous Prompt: Attack Techniques For Language Models
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs