Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs

arXiv:2608.12675 · cs.AI, cs.CR · Submitted 2026-08-13 · Read on arXiv

King Saud University · Institution of Public Administration

cs.AI, cs.CR

Submitted: 2026-08-13

Updated: 2026-09-05

Comments: Submitted to Knowledge-Based Systems Journal

Code: https://github.com/Saleh-Almohaimeed/SEAG

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper introduces the Sensitive Entity Alias Generator (SEAG), a privacy-preserving framework for Retrieval-Augmented Generation (RAG) systems.

Terminology

Summary

The paper introduces the Sensitive Entity Alias Generator (SEAG), a privacy-preserving framework for Retrieval-Augmented Generation (RAG) systems. The framework addresses a privacy issue often overlooked in RAG research: external third-party generators (such as GPT-5 or Claude-4 sonnet) have access to the user's query and retrieved documents, which may contain confidential information that could be misused. SEAG enables users to utilize powerful third-party generators without disclosing sensitive information.

The SEAG framework introduces a lightweight local model (approximately 3–4 billion parameters) that is deployed between the retriever and the generator. This model locates sensitive entities, generates corresponding aliases, and constructs an entity replacement table. The table is used to replace sensitive words in the user's query and in the retrieved documents before they are forwarded to an external generator. After the generator produces a response, the entity replacement table is used again to map any aliases back to their original entities, resulting in a final answer that preserves both the utility of the external LLM and the privacy of sensitive information.

Two datasets were constructed: one for fine-tuning SEAG models to generate entity replacement tables (640 samples from 8 domains: news, legal, medical, finance, biology, chemistry, history, and resume), and another for evaluating the entire SEAG framework (600 samples from 6 domains: economic, education, cars, legal, finance, business). The evaluation domains differ from those used during fine-tuning, except for Legal and Finance.

Three LLMs (Qwen-3, LLaMA-3.2, and Phi-4) were fine-tuned using QLoRA, a parameter-efficient fine-tuning technique. All models are of size 3B except Qwen-3, which is 4B. The framework was evaluated on two popular generators: GPT-5 and Claude-4 sonnet.

The framework is measured by two main metrics. The Privacy metric measures whether the model correctly hides sensitive information by replacing them with meaningful words. The User metric measures whether the whole SEAG framework has been implemented successfully, meaning sensitive information has been hidden and correct results have been shown to the user.

Experimental results show that all SEAG models achieved over 80% accuracy on the User metric. For the Privacy metric, the best model was LLaMA-3.2, which scored 89.67% with Claude-4 sonnet as the generator. For the User metric, the best results were achieved by the combination of LLaMA-3.2 as the SEAG model and Claude-4-sonnet as the generator, with a score of 83.67%.

Additional analysis evaluated the ability of SEAG models Qwen-3, LLaMA-3.2, and Phi-4 to hide all sensitive entities within given documents. The results show total accuracies of 77.83%, 76.73%, and 74.91%, respectively. The economic domain consistently produced the best results for all three models, with accuracy ranging from 86.01% to 87.46%.

The paper identifies two major limitations: document length (average approximately 300 words) and the number of sensitive entities (approximately 15 per document). Future research should examine longer documents, documents with substantially larger numbers of sensitive entities, additional categories of sensitive information, multilingual settings, and stronger replacement strategies.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a Local Privacy-Scrubbing Layer to Any RAG Pipeline
  • Integrate a lightweight (3–4B parameter) local model as a mandatory pre/post-processing step between retrieval and generation.

  • This model automatically detects sensitive entities (names, IDs, financial figures, medical details) in both the user query and retrieved documents, replaces them with context-preserving aliases, and reverses the mapping after generation.

  • The improved system can use any external LLM (e.g., GPT-5, Claude-4) without leaking confidential data, even if the external provider logs or misuses inputs.

  1. Enable Domain-Adaptive Alias Generation
  • Fine-tune the local alias generator on diverse domain corpora (legal, medical, finance, etc.) using QLoRA for efficiency.

  • The improved system can dynamically switch alias strategies based on the detected domain, ensuring that replacements remain semantically meaningful (e.g., replacing a patient name with Patient A in medical text, but a company name with Firm X in finance).

  • This preserves answer quality because the external LLM still receives coherent, realistic text rather than garbled placeholders.

  1. Implement a Two-Metric Quality Gate
  • Use the paper’s dual evaluation (Privacy metric for hiding completeness, User metric for end-to-end utility) as a runtime check.

  • The improved system can automatically reject or retry generations where aliases are not fully reversed or where the final answer contains leftover alias tokens, ensuring the user always sees correct, original entities.

  1. Optimize for Long-Context and High-Density Sensitive Data
  • Extend the local model’s capacity to handle documents longer than 300 words and more than 15 sensitive entities per document, using chunking and incremental alias table merging.

  • The improved system can process full research papers, legal contracts, or patient records (thousands of words, dozens of entities) without degradation, while still maintaining >80% utility.

  1. Add Multilingual and Cross-Script Alias Support
  • Train the alias generator on multilingual data (e.g., English, Chinese, Arabic) with script-aware replacement (e.g., transliteration for names, numeric formatting for dates).

  • The improved system can protect sensitive information in non-English RAG workflows, enabling safe use of global third-party LLMs across languages.

  1. Strengthen Alias Uniqueness and Reversibility
  • Replace simple word-level aliases with a cryptographic-style mapping (e.g., random but deterministic tokens per session) to prevent inference attacks from repeated queries.

  • The improved system can guarantee that even if an attacker sees multiple aliased documents, they cannot correlate aliases across sessions or reverse-engineer the original entities without the local mapping table.

  1. Provide a Fallback for Unrecognized Sensitive Entities
  • Add a rule-based and NER-based hybrid detector to catch entities the fine-tuned model misses (e.g., rare acronyms, custom identifiers).

  • The improved system can flag and redact any remaining sensitive tokens before sending to the external LLM, reducing the privacy failure rate from 10% to near zero.

What the Improved AI System Can Do:

  • Deploy any state-of-the-art external LLM for RAG tasks (question answering, summarization, report generation) while guaranteeing that no personal, corporate, or medical secrets leave the user’s local environment.

  • Maintain high answer accuracy (≥80%) across diverse domains, even with long documents and many sensitive entities.

  • Operate on consumer hardware (3–4B model) with minimal latency, making privacy-preserving RAG practical for real-time applications like legal research, clinical decision support, or financial analysis.

  • Automatically adapt to new domains and languages without retraining the external LLM, only the small local model.

  • Provide auditable privacy: every alias mapping is logged locally, so users can verify what was hidden and restored, and can revoke access at any time.

Abstract

Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models (LLMs) in answering user queries. Existing privacy research on RAG has focused on preventing unauthorized users from accessing sensitive data. However, another important problem that is often overlooked in RAG privacy research is that external generators have access to the query and the retrieved documents, which may contain confidential information that could potentially be misused or accessed for unintended purposes. In this paper, we introduce the Sensitive Entity Alias Generator (SEAG), a privacy-preserving framework that empowers users to utilize powerful third-party generators without disclosing sensitive information. SEAG introduces a lightweight model that locates sensitive entities, generates corresponding aliases, and constructs an entity replacement table. The table is used to replace sensitive words in the user's query and in the retrieved documents before they are forwarded to an external generator. For this purpose, two datasets were constructed: one for fine-tuning SEAG models to generate entity replacement tables, and another for evaluating the entire SEAG framework. The experimental results demonstrate the success of the SEAG framework. As for the User metric, which measures the ability of the model to provide a correct response to the user while hiding sensitive information from the external generator, all SEAG models achieved over 80% accuracy. Additional analysis further evaluated the ability of SEAG models Qwen-3, LLaMA-3.2, and Phi-4 to hide all sensitive entities within given documents. The results show good performance with total accuracies of 77.83%, 76.73%, and 74.91%, respectively.

Sources

Related papers