Addressing Corpus Knowledge Poisoning Attacks on RAG Using Sparse Attention

arXiv:2602.04711 · cs.IR, cs.AI · Submitted 2026-02-04 · Read on arXiv

cs.IR, cs.AI

Submitted: 2026-02-04

Updated: 2026-08-26

Code: https://github.com/sagie-dekel/Sparse-DocumentAttention-RAG

License: http://creativecommons.org/licenses/by/4.0/

The gist: Retrieval Augmented Generation (RAG) is a highly effective paradigm for keeping LLM-based responses up-to-date and reducing the likelihood of hallucinations.

Terminology

Abstract

Retrieval Augmented Generation (RAG) is a highly effective paradigm for keeping LLM-based responses up-to-date and reducing the likelihood of hallucinations. Yet, RAG was recently shown to be quite vulnerable to corpus knowledge poisoning: an attacker injects misleading documents to the corpus to steer an LLM's output to an undesired response. We argue that the standard causal attention mechanism in LLMs enables harmful cross-document interactions, specifically in cases of attacks. Accordingly, we introduce a novel defense approach for RAG: Sparse Document Attention RAG (SDAG). This is a block-sparse attention mechanism that disallows cross-attention between retrieved documents. SDAG requires a minimal inference-time change to the attention mask. We present an empirical evaluation of LLM-based question answering (QA) with a variety of attack strategies on RAG. We show that our SDAG method substantially outperforms the standard causal attention mechanism. We further demonstrate the clear merits of integrating SDAG with state-of-the-art RAG defense methods. Specifically, the integration results in performance that is statistically significantly better than the state-of-the-art.

Sources

Related papers