A Survey of Secure Retrieval-Augmented Generation
summary
The gist
Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but this access path also introduces security risks that existing work often conflates with inherent
In short
This survey systematically organizes security risks in Retrieval-Augmented Generation (RAG) using a framework called SLOT. It maps attack surfaces to defense layers and classifies security issues by their objective (CIA properties) and target, distinguishing between simple known queries and more realistic, adaptive attackers.
Key concepts
- SLOT View
- A comprehensive taxonomy for RAG security organized around four dimensions: Attack Surface (S), Defense Layer (L), Objective (O - CIA properties), and Target (T). This structure helps systematically categorize all potential security threats in RAG systems.
- Attack Surface (S) vs. Defense Layer (L)
- The paper divides the RAG pipeline into four attack surfaces, each corresponding to a defense layer. These surfaces map the flow of data—from external knowledge sources to final generation—allowing researchers to pinpoint exactly where vulnerabilities exist and where defenses are needed.
- Target (T1 vs. T2)
- Attacks are categorized by their target: T1 attacks focus on known queries, while the more realistic T2 attackers perform target-claim manipulation across many queries. This distinction highlights a major gap, as defenses are often weak against these adaptive, systemic attacks.
Terminology used across episodes
This episode discusses
- A Survey of Secure Retrieval-Augmented Generation · Paper Radio
- C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
- Do Multimodal RAG Systems Leak Data? A Comprehensive Evaluation of Membership Inference and Image Caption Retrieval Attacks
- RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
- SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems
- Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks · Paper Radio
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- A Survey on Knowledge-Oriented Retrieval-Augmented Generation
- TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
- Secure Retrieval-Augmented Generation against Poisoning Attacks
- Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
- Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
- Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks
- Addressing Corpus Knowledge Poisoning Attacks on RAG Using Sparse Attention · Paper Radio
- Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
- Memory Injection Attacks on LLM Agents via Query-Only Interaction
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
- Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
- Retrieval-Augmented Generation for Large Language Models: A Survey
The paper
A Survey of Secure Retrieval-Augmented Generation · Read on arXiv
The Hong Kong Polytechnic University
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "A Survey of Secure Retrieval-Augmented Generation".
Elias: Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but this access path also introduces security risks that existing work often conflates with inherent LLM flaws.
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So we've got this paper here called "A Survey of Secure Retrieval-Augmented Generation," and it seems like the authors are setting up a really clear framework for understanding security risks in RAG systems. What's the main idea they're pushing with this survey?
Elias: Well, the paper argues that most existing work gets tangled up, confusing security issues specific to RAG with general problems we already know about large language models themselves. They propose a new way to look at it using a taxonomy they call SLOT, which organizes everything around the attack surface, the defense layer, the objective being broken regarding CIA properties, and what exactly the attacker is trying to achieve.
Priya: That sounds like it could be very helpful for researchers trying to map out where vulnerabilities actually lie in these systems. I wonder how this structure helps distinguish between simple LLM flaws and problems that arise specifically because of how external knowledge is brought in.
Nadia: Exactly, Priya, the core claim seems to be that an attacker doesn't necessarily need to touch the model or even the prompt itself; they can cause harm by tampering with things outside the model—like changing what gets retrieved or how that retrieved context is used. That’s why this taxonomy is so crucial for understanding where we need to focus our security efforts.
Elias: They lay out a six-stage pipeline for RAG, moving from external sources to generation, and then they map four distinct attack surfaces—S1 through S4—onto corresponding defense layers, L1 through L4. This mapping helps visualize the entire flow of data and where defenses should be placed relative to the threats.
Priya: Mapping those stages onto attack surfaces is a good way to show the physical or procedural steps where an adversary can intervene, which is important for understanding measurement and privacy risks too. I'm curious about how they categorize those surfaces across the pipeline.
Nadia: They define S1 as Knowledge Poisoning, S2 as Retrieval Result Manipulation, S3 as Retrieved-Context Exploitation, and S4 as Private Knowledge Extraction; meanwhile, L1 is Integrity and Provenance, L2 is Retrieval-time Access Hardening, L3 is Post-Retrieval Isolation and Robust Generation, and L4 is Access Control and Confidentiality. That's a very concrete way to visualize the security posture of a RAG system.
Elias: And then they introduce the Objective (O) axis—Integrity, Availability, and Confidentiality—and the Target (T) axis, which differentiates between T1, attacks on known queries, and T2, which involves target-claim manipulation across a distribution of queries. That distinction is what makes their framework really robust for analyzing attack types.
Priya: The shift from T1 to T2 seems significant because it moves the focus from testing a system with specific inputs to understanding how an attacker can subtly shift the system's stance over time based on how users query it. This suggests a more realistic threat model for real-world deployment scenarios.
Paper summary: Nadia: Precisely, Priya; T2 attacks are much harder to defend against because they aren't just one isolated event; they require manipulating the system across many interactions, and this is where many current defenses seem inadequate. The paper emphasizes that most existing research still focuses on T1, which seems a real blind spot for the community.
Elias: They also point out some structural mismatches between the attacks and defenses, noting that attacks like knowledge poisoning can be persistent because malicious content can stay in a shared store, while defenses are often concentrated further downstream where the context is already being used by the LLM.
Priya: That mismatch between where the attack starts and where we put our safeguards really tells us something about current design choices; it suggests that upstream controls might need to be much stronger than what's currently implemented. What does this structural mismatch imply for developing better evaluation methods?
Nadia: It implies we need evaluation that moves beyond just checking if a single query fails, and instead needs to test the system's robustness against those adaptive T2 manipulation strategies they described in the survey. That’s a big challenge for us as applied researchers.
Elias: Looking ahead, the paper suggests several directions for future work, such as developing adaptive defenses that don't rely on blind-spot assumptions and focusing specifically on the confidentiality surface which is currently under-served by research. They also mention needing persistence-aware evaluation for systems involving multimodal or agentic RAG architectures.
Priya: Those future directions sound very practical because they target the specific weaknesses identified in the taxonomy, like building defenses that are aware of long-term persistence rather than just immediate input validation. It seems like they're pushing for a more holistic security approach across the entire pipeline.
Nadia: So, to wrap up this discussion on "A Survey of Secure Retrieval-Augmented Generation," the main contribution is providing this SLOT view—the way they organize security along the surface, layer, objective, and target axes—which gives us a unified language to talk about RAG vulnerabilities.
Elias: And another key point is defining that target-level problem definition, T1 versus T2, which really helps us understand the difference between testing a known input and modeling an attacker who can subtly manipulate claims across many queries.
Priya: I think what this whole survey really contributes to the broader field is making sure we aren't just focusing on one narrow aspect of RAG security when there are so many interconnected pathways for harm. It sets a much clearer roadmap for where future research should go, especially concerning those structural mismatches they pointed out.
Paper summary: Nadia: It does give us that roadmap by clearly showing the pipeline and how each stage relates to a specific defense layer, which helps us decide exactly where to invest our security efforts first when we're building these systems.
Elias: And I think the implication for cryptography is that understanding S4, Private Knowledge Extraction, forces us to consider what happens when retrieval channels are used not just for information retrieval but as a means to covertly exfiltrate protected data back out through the generation process.
Priya: That's a good point, Elias; if we can map those risks clearly, we can start designing better access control mechanisms that specifically target that extraction surface without having to over-engineer the whole system unnecessarily.
Nadia: So, in short, this paper gives us the comprehensive taxonomy for RAG security by organizing it systematically around four key axes—S, L, O, and T—which helps us see exactly where the gaps are between what's being attacked and what's being defended against.
Elias: And because they highlight that gap between fluent attacks and concentrated defenses, the implication is that we need to move our defensive thinking upstream toward those initial control points where persistence begins.
Priya: That structural mismatch is a very telling observation; it tells us that simply adding more filters downstream isn't going to solve the problem of persistent knowledge poisoning or retrieval manipulation.
Nadia: Exactly, Priya; we need defenses that are built into the ingestion and indexing stages themselves, rather than just relying on post-retrieval checks when the model is already processing potentially compromised context.
Elias: And for those who are interested in the cryptographic side, their work on T2 suggests that any security proof must account for adversaries who can perform target-claim manipulation across a distribution of inputs, not just a single fixed query.
Priya: That makes sense; if we're designing protocols or systems, we have to consider the statistical properties of the attack space defined by that T2 attacker, which is much more complex than just assuming a single input vector.
Nadia: So, to conclude on this paper's impact, "A Survey of Secure Retrieval-Augmented Generation" provides the essential framework for researchers to systematically categorize and address RAG security risks based on a pipeline view and a detailed taxonomy called SLOT.
Elias: The implication is that we can finally start moving past conflating RAG security problems with inherent LLM flaws by having a concrete, organized structure to analyze the specific points of failure in the retrieval-augmentation process.
Priya: It really sets the stage for much more targeted evaluation, moving away from general benchmarks toward metrics that specifically test resilience against those defined attack surfaces and objective breaches.
Nadia: That’s right; it gives us a shared vocabulary to discuss security risks in RAG, which is a huge step forward in making this area of applied security research more coherent and actionable for everyone involved.
Conclusion: Nadia: So, we've just finished looking at how this paper structures its argument around these four axes—the surface, layer, objective, and target—which really gives us a clear map of where RAG security actually lives.
Elias: I agree with that summary; the way they break down the attack surfaces into S1 through S4 provides a solid foundation for seeing exactly what kind of tampering is possible.
Priya: From my side, seeing those objectives like integrity and confidentiality clearly laid out helps me understand what kind of privacy risks we're really looking at when we talk about these systems.
Nadia: Exactly, and the paper’s focus on distinguishing between T1 and T2 attackers seems to be a really smart way to frame the problem in a more realistic way for deployment scenarios.
Elias: That distinction is vital because modeling those target-claim manipulations across a query distribution is where you'll find the most interesting cryptographic assumptions to test.
Priya: I wonder if this framework will eventually allow us to develop metrics that actually measure the privacy leakage happening at surface S4, private knowledge extraction.
Nadia: That’s exactly what I mean; we need those evaluation methods that move beyond just checking a single query's output and test resilience across these defined attack surfaces.
Elias: And if the authors manage to map those defense layers L1 through L4 effectively, it gives us concrete targets for designing stronger access controls upstream.
Priya: It’s exciting because this survey essentially tells us where the current blind spots are, which is a huge step toward building better defenses for privacy concerns.
Nadia: Indeed, and I think the implications here are that we can finally start talking about RAG security in a way that doesn't get lost in general LLM discussion.
Elias: The paper’s authors have done a good job of connecting the theoretical attack vectors to practical pipeline stages, which is something I appreciate as a cryptographer.
Priya: So, we’re looking at how this taxonomy might change the way privacy researchers approach measuring data exposure in these complex architectures.
Nadia: Exactly; this structure gives us a shared vocabulary to discuss RAG security risks without getting bogged down in too much noise about the underlying AI models themselves.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel