A Survey of Secure Retrieval-Augmented Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "A Survey of Secure Retrieval-Augmented Generation".
Elias: Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but this access path also introduces security risks that existing work often conflates with inherent LLM flaws.
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So we've got this paper here called "A Survey of Secure Retrieval-Augmented Generation," and it seems like the authors are setting up a really clear framework for understanding security risks in RAG systems. What's the main idea they're pushing with this survey?
Elias: Well, the paper argues that most existing work gets tangled up, confusing security issues specific to RAG with general problems we already know about large language models themselves. They propose a new way to look at it using a taxonomy they call SLOT, which organizes everything around the attack surface, the defense layer, the objective being broken regarding CIA properties, and what exactly the attacker is trying to achieve.
Priya: That sounds like it could be very helpful for researchers trying to map out where vulnerabilities actually lie in these systems. I wonder how this structure helps distinguish between simple LLM flaws and problems that arise specifically because of how external knowledge is brought in.
Nadia: Exactly, Priya, the core claim seems to be that an attacker doesn't necessarily need to touch the model or even the prompt itself; they can cause harm by tampering with things outside the model—like changing what gets retrieved or how that retrieved context is used. That’s why this taxonomy is so crucial for understanding where we need to focus our security efforts.
Elias: They lay out a six-stage pipeline for RAG, moving from external sources to generation, and then they map four distinct attack surfaces—S1 through S4—onto corresponding defense layers, L1 through L4. This mapping helps visualize the entire flow of data and where defenses should be placed relative to the threats.
Priya: Mapping those stages onto attack surfaces is a good way to show the physical or procedural steps where an adversary can intervene, which is important for understanding measurement and privacy risks too. I'm curious about how they categorize those surfaces across the pipeline.
Nadia: They define S1 as Knowledge Poisoning, S2 as Retrieval Result Manipulation, S3 as Retrieved-Context Exploitation, and S4 as Private Knowledge Extraction; meanwhile, L1 is Integrity and Provenance, L2 is Retrieval-time Access Hardening, L3 is Post-Retrieval Isolation and Robust Generation, and L4 is Access Control and Confidentiality. That's a very concrete way to visualize the security posture of a RAG system.
Elias: And then they introduce the Objective (O) axis—Integrity, Availability, and Confidentiality—and the Target (T) axis, which differentiates between T1, attacks on known queries, and T2, which involves target-claim manipulation across a distribution of queries. That distinction is what makes their framework really robust for analyzing attack types.
Priya: The shift from T1 to T2 seems significant because it moves the focus from testing a system with specific inputs to understanding how an attacker can subtly shift the system's stance over time based on how users query it. This suggests a more realistic threat model for real-world deployment scenarios.
Paper summary: Nadia: Precisely, Priya; T2 attacks are much harder to defend against because they aren't just one isolated event; they require manipulating the system across many interactions, and this is where many current defenses seem inadequate. The paper emphasizes that most existing research still focuses on T1, which seems a real blind spot for the community.
Elias: They also point out some structural mismatches between the attacks and defenses, noting that attacks like knowledge poisoning can be persistent because malicious content can stay in a shared store, while defenses are often concentrated further downstream where the context is already being used by the LLM.
Priya: That mismatch between where the attack starts and where we put our safeguards really tells us something about current design choices; it suggests that upstream controls might need to be much stronger than what's currently implemented. What does this structural mismatch imply for developing better evaluation methods?
Nadia: It implies we need evaluation that moves beyond just checking if a single query fails, and instead needs to test the system's robustness against those adaptive T2 manipulation strategies they described in the survey. That’s a big challenge for us as applied researchers.
Elias: Looking ahead, the paper suggests several directions for future work, such as developing adaptive defenses that don't rely on blind-spot assumptions and focusing specifically on the confidentiality surface which is currently under-served by research. They also mention needing persistence-aware evaluation for systems involving multimodal or agentic RAG architectures.
Priya: Those future directions sound very practical because they target the specific weaknesses identified in the taxonomy, like building defenses that are aware of long-term persistence rather than just immediate input validation. It seems like they're pushing for a more holistic security approach across the entire pipeline.
Nadia: So, to wrap up this discussion on "A Survey of Secure Retrieval-Augmented Generation," the main contribution is providing this SLOT view—the way they organize security along the surface, layer, objective, and target axes—which gives us a unified language to talk about RAG vulnerabilities.
Elias: And another key point is defining that target-level problem definition, T1 versus T2, which really helps us understand the difference between testing a known input and modeling an attacker who can subtly manipulate claims across many queries.
Priya: I think what this whole survey really contributes to the broader field is making sure we aren't just focusing on one narrow aspect of RAG security when there are so many interconnected pathways for harm. It sets a much clearer roadmap for where future research should go, especially concerning those structural mismatches they pointed out.
Paper summary: Nadia: It does give us that roadmap by clearly showing the pipeline and how each stage relates to a specific defense layer, which helps us decide exactly where to invest our security efforts first when we're building these systems.
Elias: And I think the implication for cryptography is that understanding S4, Private Knowledge Extraction, forces us to consider what happens when retrieval channels are used not just for information retrieval but as a means to covertly exfiltrate protected data back out through the generation process.
Priya: That's a good point, Elias; if we can map those risks clearly, we can start designing better access control mechanisms that specifically target that extraction surface without having to over-engineer the whole system unnecessarily.
Nadia: So, in short, this paper gives us the comprehensive taxonomy for RAG security by organizing it systematically around four key axes—S, L, O, and T—which helps us see exactly where the gaps are between what's being attacked and what's being defended against.
Elias: And because they highlight that gap between fluent attacks and concentrated defenses, the implication is that we need to move our defensive thinking upstream toward those initial control points where persistence begins.
Priya: That structural mismatch is a very telling observation; it tells us that simply adding more filters downstream isn't going to solve the problem of persistent knowledge poisoning or retrieval manipulation.
Nadia: Exactly, Priya; we need defenses that are built into the ingestion and indexing stages themselves, rather than just relying on post-retrieval checks when the model is already processing potentially compromised context.
Elias: And for those who are interested in the cryptographic side, their work on T2 suggests that any security proof must account for adversaries who can perform target-claim manipulation across a distribution of inputs, not just a single fixed query.
Priya: That makes sense; if we're designing protocols or systems, we have to consider the statistical properties of the attack space defined by that T2 attacker, which is much more complex than just assuming a single input vector.
Nadia: So, to conclude on this paper's impact, "A Survey of Secure Retrieval-Augmented Generation" provides the essential framework for researchers to systematically categorize and address RAG security risks based on a pipeline view and a detailed taxonomy called SLOT.
Elias: The implication is that we can finally start moving past conflating RAG security problems with inherent LLM flaws by having a concrete, organized structure to analyze the specific points of failure in the retrieval-augmentation process.
Priya: It really sets the stage for much more targeted evaluation, moving away from general benchmarks toward metrics that specifically test resilience against those defined attack surfaces and objective breaches.
Nadia: That’s right; it gives us a shared vocabulary to discuss security risks in RAG, which is a huge step forward in making this area of applied security research more coherent and actionable for everyone involved.
Conclusion: Nadia: So, we've just finished looking at how this paper structures its argument around these four axes—the surface, layer, objective, and target—which really gives us a clear map of where RAG security actually lives.
Elias: I agree with that summary; the way they break down the attack surfaces into S1 through S4 provides a solid foundation for seeing exactly what kind of tampering is possible.
Priya: From my side, seeing those objectives like integrity and confidentiality clearly laid out helps me understand what kind of privacy risks we're really looking at when we talk about these systems.
Nadia: Exactly, and the paper’s focus on distinguishing between T1 and T2 attackers seems to be a really smart way to frame the problem in a more realistic way for deployment scenarios.
Elias: That distinction is vital because modeling those target-claim manipulations across a query distribution is where you'll find the most interesting cryptographic assumptions to test.
Priya: I wonder if this framework will eventually allow us to develop metrics that actually measure the privacy leakage happening at surface S4, private knowledge extraction.
Nadia: That’s exactly what I mean; we need those evaluation methods that move beyond just checking a single query's output and test resilience across these defined attack surfaces.
Elias: And if the authors manage to map those defense layers L1 through L4 effectively, it gives us concrete targets for designing stronger access controls upstream.
Priya: It’s exciting because this survey essentially tells us where the current blind spots are, which is a huge step toward building better defenses for privacy concerns.
Nadia: Indeed, and I think the implications here are that we can finally start talking about RAG security in a way that doesn't get lost in general LLM discussion.
Elias: The paper’s authors have done a good job of connecting the theoretical attack vectors to practical pipeline stages, which is something I appreciate as a cryptographer.
Priya: So, we’re looking at how this taxonomy might change the way privacy researchers approach measuring data exposure in these complex architectures.
Nadia: Exactly; this structure gives us a shared vocabulary to discuss RAG security risks without getting bogged down in too much noise about the underlying AI models themselves.
The Hong Kong Polytechnic University
cs.CR, cs.AI
Submitted: 2026-04-09
Updated: 2026-10-07
Comments: Accepted at EMNLP 2026
Code: https://github.com/TreeAI-Lab/Awesome-RAG-Security
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 83/100
The gist: Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but this access path also introduces security risks that existing work often conflates with inherent
Key concepts
- SLOT View
- A comprehensive taxonomy for RAG security organized around four dimensions: Attack Surface (S), Defense Layer (L), Objective (O - CIA properties), and Target (T). This structure helps systematically categorize all potential security threats in RAG systems.
- Attack Surface (S) vs. Defense Layer (L)
- The paper divides the RAG pipeline into four attack surfaces, each corresponding to a defense layer. These surfaces map the flow of data—from external knowledge sources to final generation—allowing researchers to pinpoint exactly where vulnerabilities exist and where defenses are needed.
- Target (T1 vs. T2)
- Attacks are categorized by their target: T1 attacks focus on known queries, while the more realistic T2 attackers perform target-claim manipulation across many queries. This distinction highlights a major gap, as defenses are often weak against these adaptive, systemic attacks.
Terminology
Summary
Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but this access path also introduces security risks that existing work often conflates with inherent LLM flaws. The central problem is that an attacker need not touch the model or the user-visible prompt; by tampering with external content, retrieval process, or disclosure behavior, it can drive harmful evidence into the generator or read protected knowledge back out. This paper presents a comprehensive and systematic survey dedicated to RAG security organized under a single explicit and hierarchical taxonomy called SLOT.
The gist
The SLOT view organizes RAG security along one backbone—the attack Surface (S), mirrored by the defense Layer (L)—and two cross-cutting axes: the Objective (O) it breaks following the CIA properties, and the Target (T) it pursues, from a single known query (T1) to target-claim manipulation across a query distribution (T2).
RAG Pipeline and Security Surfaces
The survey views RAG processes through an external-knowledge access pipeline divided into six stages: external sources with raw content; ingestion and indexing that turn raw content into chunks and embeddings; retriever and reranking select candidate evidence; context assembly formats the selected evidence into the model-visible prompt; generation produces the answer; and response and remediation. This pipeline is mapped to four attack surfaces (S1-S4) mirroring defense layers (L1-L4):
(S1: Knowledge Poisoning)
(S2: Retrieval Result Manipulation)
(S3: Retrieved-Context Exploitation)
(S4: Private Knowledge Extraction)
Attack Objectives and Targets
The objective (O) is classified by the classical CIA triad [76]: Integrity (O1), Availability (O2), and Confidentiality (O3). The target (T) specifies the attacker’s intent. A T1 attacker targets known queries, while a T2 attacker performs target-claim manipulation: fixing a claim about an entity and shifting the system’s stance across queries touching it. This distinction is crucial because T2 is more realistic and harder than T1, yet attacks are few and defenses nearly absent against it, creating a structural blind spot.
Defense Layers and Structural Mismatches
Defenses mirror the attack surfaces: L1 (Integrity & Provenance), L2 (Retrieval-time Access Hardening), L3 (Post-Retrieval Isolation & Robust Generation), and L4 (Access Control & Confidentiality). The survey exposes two structural mismatches: attacks are increasingly fluent and corpus-fitting, while defenses remain concentrated downstream and comparatively thin at the upstream control points where persistent harm begins.
Key Observations on Mismatches
The SLOT map highlights that knowledge poisoning (S1) is persistent because malicious content can stay in a shared store, while retrieval result manipulation (S2) changes what evidence reaches the model. Retrieved-context exploitation (S3) changes how the model follows retrieved content after the downstream boundary, and private knowledge extraction (S4) turns retrieval into a channel for recovering protected knowledge. Across these surfaces, integrity attacks (O1) are the most crowded, availability attacks (O2) center on refusal or failure, and confidentiality attacks (O3) mainly appear as privacy and extraction risks. Most work still operates at T1, with only a small but growing line of attacks reaching the T2 target across S1, S2, and S4.
Future Directions
The paper suggests five directions for advancing secure RAG: realistic attacker targets (I1), no-blind-spot adaptive defense (I2–I3), the under-served confidentiality surface (I4), and persistence-aware evaluation for multimodal and agentic RAG (I5). Future work should specify threats by a target claim with the victim’s query distribution, move defenses from per-query fact-checking toward claim-level auditing, build a no-blind-spot wall along the pipeline, focus on selective disclosure for private corpora, and prepare evaluation for multimodal and agentic RAG systems.
Summary of Contributions
The primary contributions are:
-
The SLOT view: organizing RAG security along one backbone (S↔L) and two cross-cutting axes (O and T).
-
A target-level problem definition (T): distinguishing T1 from the realistic T2 attacker.
-
An objective level with a principled basis (O): classifying attacks by CIA properties [76].
-
A pipeline backbone (S & L): defining four attack surfaces and mirroring them with defense layers along the knowledge-access pipeline.
-
A diagnosis of structural mismatch: exposing field-level gaps where attacks are fluent and defenses are concentrated downstream, while also noting the scarcity of evaluation against adaptive attackers.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed this comprehensive taxonomy of Retrieval-Augmented Generation (RAG) security. The core insight is that securing RAG requires moving beyond treating LLM flaws as the primary threat and instead framing it as securing the external knowledge access path across a structured pipeline (S1-S4).
Based on the findings, here are specific, actionable improvements for AI systems:
The improved AI system will transition from a reactive answer generator
to a proactive, security-aware knowledge orchestrator.
The key capabilities gained are:
-
Predictive defense against upstream knowledge corruption (S1).
-
Contextual resilience against downstream exploitation (S3).
-
Claim-level stance shifting across complex query distributions (T2 defense).
Here are the specific improvements:
-
The system must implement a multi-layered defense strategy mirroring the SLOT taxonomy, specifically focusing on strengthening Layer 1 (Knowledge-Base Integrity) and Layer 2 (Retrieval Hardening), as these are identified as structurally weak points against persistent upstream attacks.
-
Implement robust provenance tracking for all ingested knowledge, utilizing cryptographic methods (like D-RAG or ProofCarrying Answers) to verify the origin and integrity of every chunk before it enters the index, directly countering Knowledge Poisoning (S1).
-
Integrate consistency-based aggregation mechanisms during retrieval to mitigate Retrieval Result Manipulation (S2). This means the system will not rely on a single retrieved document but will calculate reliability scores based on source credibility or evidence majority to filter out adversarial ranking shifts.
-
Deploy sophisticated post-retrieval isolation techniques (L3) that use attention variance signals or lightweight ML filters to detect when retrieved context is being used as an indirect instruction carrier (S3), allowing the system to block harmful behaviors before they are executed by the generator.
-
Develop a T2-specific defense mechanism: Instead of focusing solely on fixing answers for known queries (T1), the system will be designed to audit and stabilize its
stance
across an entire distribution of related future queries. This involves training or fine-tuning models with robustness against counterfactual context, ensuring the system maintains a consistent, safe stance on specific claims even when faced with unknown or adversarial query patterns. -
In high-stakes domains (e.g., medical, financial), implement Layer 4 controls that use selective disclosure and differential privacy protocols to prevent Private Knowledge Extraction (S4) by reducing what can be inferred from the retrieved corpus, moving beyond simple prompt-level refusal.
Sources
- C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
- Do Multimodal RAG Systems Leak Data? A Comprehensive Evaluation of Membership Inference and Image Caption Retrieval Attacks
- RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
- SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems
- Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- A Survey on Knowledge-Oriented Retrieval-Augmented Generation
- TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
- Secure Retrieval-Augmented Generation against Poisoning Attacks
- Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
- Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
- Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks
- Addressing Corpus Knowledge Poisoning Attacks on RAG Using Sparse Attention
- Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
- Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
- Retrieval-Augmented Generation for Large Language Models: A Survey
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs