RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
Pushkal Kumar, Tucker Nielson, Tanish Kolhe, Shubham Zala, Vincent Li
cs.LG, cs.CR, cs.IR
Submitted: 2026-07-28
Comments: Accepted to NeurIPS ResponsibleFM 2025, AAAI FrontierIR 2026
Code: https://github.com/RAGuard-AI/RAGuard
Project page: https://christophm.github.io/interpretable-ml-book
License: http://creativecommons.org/licenses/by/4.0/
The gist: Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate
Terminology
Abstract
Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate retrieved evidence. We introduce RAGuard, a layered defense against factual corpus-poisoning attacks on RAG pipelines. The first layer adversarially fine-tunes a dense retriever on synthetic poisoned documents (fabricated facts, contradictions, and reasoning traps), teaching it to downrank malicious passages before generation. The second layer, the Zero-Knowledge Inference Patch ZKIP, is a label-free, black-box filter: for each retrieved document, it performs a leave-one-out decode and scores the document by the semantic shift and output-entropy change that its removal induces. ZKIP requires no poison labels, no ground-truth answers, and no access to model internals; it compares the model's own answers under counterfactual contexts. On poisoned Natural Questions at 5--30% poison ratios, adversarial retriever training alone reduces but does not eliminate attack success, while ZKIP drives the measured attack success rate to 0.000 in every defended configuration, keeping Recall@5 within 0.03 of the clean-corpus baseline. Supervised analyses on both Natural Questions and BEIR (NFCorpus) confirm that the counterfactual signals ZKIP relies on carry learnable poison structure. The defense costs k + 1 generator passes per query (6 times for k = 5); we analyze batching and early-stopping approximations that reduce this overhead. We also show that keyword-preserving poisons leave lexical retrievers such as BM25 essentially unaffected, an observation that delineates the boundary of the threat model. Code, datasets, and evaluation harnesses are released for reproducibility.
Sources
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- Unsupervised Dense Information Retrieval with Contrastive Learning
- Atlas: Few-shot Learning with Retrieval Augmented Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Backdoor Attacks on Dense Retrieval via Public and Unintentional Triggers
- Black-box Backdoor Defense via Zero-shot Image Purification
- Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Corpus Poisoning via Approximate Greedy Gradient Descent
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks