Addressing Corpus Knowledge Poisoning Attacks on RAG Using Sparse Attention
cs.IR, cs.AI
Submitted: 2026-02-04
Updated: 2026-08-26
Code: https://github.com/sagie-dekel/Sparse-DocumentAttention-RAG
License: http://creativecommons.org/licenses/by/4.0/
The gist: Retrieval Augmented Generation (RAG) is a highly effective paradigm for keeping LLM-based responses up-to-date and reducing the likelihood of hallucinations.
Terminology
Abstract
Retrieval Augmented Generation (RAG) is a highly effective paradigm for keeping LLM-based responses up-to-date and reducing the likelihood of hallucinations. Yet, RAG was recently shown to be quite vulnerable to corpus knowledge poisoning: an attacker injects misleading documents to the corpus to steer an LLM's output to an undesired response. We argue that the standard causal attention mechanism in LLMs enables harmful cross-document interactions, specifically in cases of attacks. Accordingly, we introduce a novel defense approach for RAG: Sparse Document Attention RAG (SDAG). This is a block-sparse attention mechanism that disallows cross-attention between retrieved documents. SDAG requires a minimal inference-time change to the attention mask. We present an empirical evaluation of LLM-based question answering (QA) with a variety of attack strategies on RAG. We show that our SDAG method substantially outperforms the standard causal attention mechanism. We further demonstrate the clear merits of integrating SDAG with state-of-the-art RAG defense methods. Specifically, the integration results in performance that is statistically significantly better than the state-of-the-art.
Sources
- DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
- The Llama 3 Herd of Models
- Unsupervised Dense Information Retrieval with Contrastive Learning
- Mistral 7B
- A Survey of Large Language Models Attribution
- Large Language Models: A Survey
- GPT-4o System Card
- RIPRAG: Hack a Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning
- BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
- Qwen2 Technical Report
- TrustRAG: Enhancing Robustness and Trustworthiness in Retrieval-Augmented Generation
- Qwen3 Technical Report
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG