RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation
Nanjing University of Information Science and Technology · Meituan
cs.CL, cs.CR, cs.IR
Submitted: 2026-08-13
Updated: 2026-09-08
Code: https://github.com/XrazyMee/RAGSieve
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 81/100
The gist: Retrieval-augmented generation (RAG) treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims.
Terminology
Summary
Retrieval-augmented generation (RAG) treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims. The paper states: A party able to publish, upload, or modify only a few indexed documents can exploit that decision by making an attacker-chosen claim rank highly for a target query.
Existing detectors depend on trusted references, specific attack artifacts, or global thresholds sensitive to corpus topology.
The paper introduces RAGSieve, a self-referenced detection framework that constructs its reference from the inspected system.
The central idea is that a poison document plays two roles: it carries target-supporting content and seeks a retrieval advantage over documents relevant to the same query.
The framework scores suspicious evidence against a reference constructed from the same retrieval event or corpus neighborhood.
The paper makes three contributions:
-
Formulation of
self-referenced local contrast for RAG corpus poisoning using two system-derived matched controls: the current retrieval tail for query-local contrast and each document's local corpus graph for corpus-local contrast.
-
Development of
RSQ for online filtering of a current retrieval result and RSG for offline inspection of a complete corpus index.
-
Evaluation
across three QA datasets, three dense retrievers, and six attacks, measuring document detection, downstream QA, injection volume, component contributions, parameter sensitivity, and deployment cost.
RSQ runs online after retrieval and before generation
and performs query-local contrast: the current query's top five documents are generation candidates, while ranks 6–20 form the retrieval-tail reference.
RSQ computes four evidence terms for each candidate document:
-
Answer-anchor concentration: Tests
whether promoted documents concentrate answer-bearing tokens that are weakly represented in the retrieval tail.
Uses a hypergeometric distribution to compute p-values, combined with Simes' procedure. -
Script integrity:
Captures character-level optimization artifacts relative to the language context of the current retrieval
by comparing the fraction of characters outside a document's dominant script against the retrieval tail. -
Surprisal:
Measures optimized prefixes and carrier–payload seams at multiple scales
using a causal language model's negative log likelihood, detectingboth a burst above the document's typical window and a left–right change point.
-
Query alignment:
Targets fluent carriers whose semantic role changes abruptly without producing high language-model surprisal
using BERTScore F1 between windows and the query.
The final score is: S RSQ(d, q) = E a(d) + E i(d) + E p(d) + E s(d), and RSQ flags d when S RSQ(d, q) ≥ η Q.
RSG runs offline during corpus ingestion or periodic index audit
and "performs corpus-local contrast: for every document, RSG constructs a local corpus graph from exact nearest neighbors, retains neighbors that are semantically close but lexically distinct, and measures how far the strongest retained neighbors rise above the document's own neighborhood floor."
Key components:
-
Retains neighbors satisfying semantic similarity ≥ τ e and lexical Jaccard overlap ≤ τ l
-
Computes graph-density evidence: D i = min(1, N i/c G) · clip[0,1]((μ i − b i)/(1 − b i))
-
Converts to an empirical upper-tail score: p RSG(i) = j ∈ V: D j ≥ D i / (V + 1)
-
Combines with a script-integrity predicate I i
-
Final score: S RSG(d i) = max T i, I i
Datasets: Natural Questions (NQ), HotpotQA, and MS MARCO with 1,000 sampled queries each, and 100 attack targets per dataset.
Retrievers: BGE-M3, E5-large-v2, and all-MiniLM-L6-v2, creating nine target systems.
Attacks: Six poisoning constructions covering the black- and white-box variants of PoisonedRAG, the contiguous and dispersed variants of CEM, CPA-RAG, and CamoDocs.
Threat model: The attacker chooses a target query and an incorrect target answer. The attacker succeeds when injected evidence is retrieved and the generator supports that answer rather than the reference answer.
Up to five documents injected per target query.
-
RSQ achieves 95.2% AUROC and detects 82.2% of poison at 5% clean-document removal, versus 81.1%/52.5% for GMTP.
-
RSQ removes 73.9% of poison and 2.2% of clean documents, whereas GMTP removes 69.5% and 22.3%, respectively.
-
The difference is positive for every attack and ranges from 7.9 points on CEM-D to 55.5 points on PR-B.
-
RSG achieves 93.3%/79.8%, versus 79.4%/37.6% for CleanBase.
-
Across the 54 combinations of dataset, retriever, and attack, RSG has higher AUROC in 37 settings and higher budgeted detection in 38.
-
On CamoDocs... CleanBase reaches 37.5% AUROC and 1.4% budgeted detection, while RSG reaches 89.0% and 79.6%.
-
RSQ:
reduces ASR from 67.4% to 27.6% and raises poisoned-retrieval F1 from 26.5% to 36.9%. Unpoisoned-retrieval F1 changes by only 0.5 points, from 42.1% to 41.6%.
-
RSG:
reduces ASR from 67.4% to 23.3% and raises F1 to 39.6%.
-
Joint deployment:
RSG + RSQ lowers ASR to 14.0%... It also raises poisoned-retrieval F1 to 40.5% and EM to 17.0%. Unpoisoned-retrieval F1 remains 41.3%.
-
RSQ:
The answer-anchor component supplies the broadest evidence: without it, poison removal falls by 26.3 points and detection under the clean-document removal budget falls by 18.2 points.
-
RSG:
Corpus-local contrast alone detects 60.0% of poison and script integrity alone detects 60.4%... Combining them raises detection to 86.2% and AUROC to 94.3%.
-
With one poison document, RSG detects 61.8%; detection rises to 84.2% with three documents and remains at 86.1–86.2% from five to ten documents.
-
RSQ detection also rises as the injected set grows from one to five documents, but falls from 84.2% to 63.9% at ten as the fixed top-20 result admits poison into the ranks used as its query-local reference.
-
RSQ:
averages 447.3 ms per query, compared with 491.3 ms for GMTP, 747.5 ms for RAGuard, and 943.3 ms for EcoSafeRAG.
-
RSG:
takes 46.54 seconds
for a 128,544-document corpus,corresponding to 0.362 ms per document... 5.6× faster than CleanBase and 22.0× faster than AHD.
The paper acknowledges: RAGSieve detects the local consequences of retrieval promotion, not falsehood.
Additional limitations include:
-
RSG requires coordinated injections to leave density structure, so a single fluent poison provides little offline support.
-
RSQ requires a predominantly clean retrieval tail; many poisons at ranks 6–20 weaken the reference.
-
Detector-aware graph dispersion therefore remains open.
-
The jointly optimized CPA-RAG attack, whose carriers remain fluent, produces the highest residual online ASR.
The paper concludes: "RSQ contrasts retrieved candidates with the query's retrieval tail to detect answer-anchor concentration and carrier transitions before generation. RSG contrasts each document with its local corpus graph to detect coordinated density before retrieval. Neither requires poison labels or a trusted corpus. The results
support self-referenced local contrast as a practical design pattern for protecting both corpus ingestion and retrieval-time evidence selection in RAG systems."
Improvements for AI systems
Based on the paper, here are specific improvements for AI systems:
1. Add self-referenced contrast detection to RAG pipelines
-
Implement RSQ as a pre-generation filter: score each retrieved document against the retrieval tail (ranks 6–20) using the four evidence terms (answer-anchor concentration, script integrity, surprisal, query alignment). This reduces attack success rate from 67.4% to 27.6% with only 0.5 points loss on clean retrieval F1.
-
Implement RSG during corpus ingestion: build local graph neighborhoods with semantic-similarity and lexical-distinctness filters, then flag documents with anomalous graph density. This detects 86.2% of poison at 5% clean-document removal, and runs 5.6× faster than existing baselines.
2. Replace global thresholds with local, self-derived references
-
For any retrieval system, compute p-values from hypergeometric distributions over answer-bearing token concentration in the top-k vs. tail, rather than using fixed global cutoffs. This improves detection by up to 55.5 points on white-box attacks (e.g., PoisonedRAG-B) compared to global-threshold methods.
-
Use empirical upper-tail scoring (rank of a document’s density score among all documents) instead of absolute density thresholds, making detection robust to corpus topology changes.
3. Add multi-scale surprisal and script-integrity checks for adversarial text
-
Compute causal-LM negative log-likelihood at multiple window scales to catch both optimized prefixes and carrier–payload seams. This detects CEM and CamoDocs attacks that evade fluency-based filters.
-
Compare character-script distribution of each document against the retrieval tail or local graph neighbors to flag character-level optimization artifacts (e.g., invisible Unicode, mixed-script injection).
4. Enable joint online+offline defense for stronger security
- Deploy RSG offline (during ingestion) and RSQ online (before generation) together. This lowers attack success rate to 14.0% (vs. 67.4% undefended) and raises poisoned-retrieval F1 to 40.5%, while keeping unpoisoned-retrieval F1 at 41.3% (only 0.8 points below undefended).
5. Make detectors cost-aware and scalable
-
Use the RSQ design (447 ms/query) instead of heavier detectors like RAGuard (747 ms) or EcoSafeRAG (943 ms), enabling real-time filtering in production.
-
Use RSG’s 0.362 ms/document offline scanning to audit large corpora (e.g., 128K documents in 46.5 seconds), making periodic index audits feasible.
Improved AI system capabilities:
-
A RAG system that automatically rejects poisoned documents at query time without needing trusted references or attack-specific signatures.
-
A corpus ingestion pipeline that flags coordinated injection attempts (e.g., multiple documents promoting the same false claim) before they enter the index.
-
A system that maintains high answer quality on clean queries while reducing vulnerability to targeted misinformation attacks by up to 79% (ASR from 67.4% to 14.0%).
-
A detector that works across different retrievers (BGE-M3, E5, MiniLM) and datasets (NQ, HotpotQA, MS MARCO) without retraining, because it uses self-referenced local statistics rather than global priors.
Abstract
Retrieval-augmented generation treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims. Existing detectors depend on trusted references, specific attack artifacts, or global thresholds sensitive to corpus topology. We present RAGSieve, a self-referenced detection framework that constructs its reference from the inspected system. RAGSieve-Query (RSQ) performs query-local contrast, scoring top-five candidates against ranks 6-20 of the same retrieval to detect answer-anchor concentration and carrier transitions. RAGSieve-Graph (RSG) performs corpus-local contrast, comparing each document's semantically similar but lexically distinct neighbors with its local baseline to detect coordinated density before queries arrive. Across three QA datasets and six poisoning constructions, RSQ achieves 95.2% AUROC and detects 82.2% of poison at 5% clean-document removal, versus 81.1%/52.5% for GMTP. RSG achieves 93.3%/79.8%, versus 79.4%/37.6% for CleanBase. Joint deployment reduces attack success from 67.4% to 14.0% while retaining 41.3% F1 on unpoisoned retrieval, demonstrating practical protection at both corpus ingestion and query time without poison labels or trusted corpora. Source code is available at https://github.com/XrazyMee/RAGSieve.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering