SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling".
Jane: The paper was written by Shixuan Sun, Siyuan Liang, Jianjie Huang, Jingzhi Li and Xiaochun Cao from Sun Yat-sen University and Nanyang Technological University and University of Chinese Academy of Sciences and Zhongguancun Academy and Institute of Information Engineering, Chinese Academy of Sciences.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everyone. Today we are digging into a paper that just hit arXiv, and the title alone got me hooked: "SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling." Jane, when you first saw that title, what went through your head?
Jane: Tom, I laughed, honestly. "Who said that?" is such a perfect question for this problem. Because with these big language models that can pull from the internet, from a database, or from their own training, you genuinely don't know where the answer came from. The paper is trying to figure out who's actually talking.
Tom: And that's the crux of it, right? We've got these Retrieval-Augmented Generation systems, RAG for short. They're supposed to make models smarter by letting them look up fresh information. But this paper says, hey, that convenience comes with a privacy cost.
Jane: Exactly. And I love that they've given this a name: source-aware membership auditing. It's not just asking "was this data in your training set?" It's asking "did this answer come from your training, or did you just pull it from that external database you're hooked up to?"
Lu: And that distinction matters more than people realize. I'm Lu, by the way, for anyone just tuning in. Think about a corporate setting. A model might be connected to a private document store. If it leaks a sentence from a confidential memo, you need to know whether that memo was in the model's original training data or whether it was retrieved live from the company's server. The fix is completely different.
Meng: Right, and from my side of things, the engineering angle, that's a huge deal. If it's in the training data, you're talking about re-training or fine-tuning to remove it. But if it's just sitting in a retrieval database, you can delete one entry and the problem is solved. That's a massive difference in cost and effort.
Tom: So the paper is basically building a lie detector for these systems, figuring out which source is actually responsible for the words coming out.
Jane: And they're doing it in a "semi-black-box" setting, which just means they can't peek inside the model's brain, but they can flip the retrieval switch on and off. That's the "controlling" part in the title.
Lu: It's a clever setup. They assume you can ask the model a question with the retrieval system on, and then ask the exact same question with it off. By comparing those two answers, you can start to isolate what the model knew on its own versus what it needed to look up.
Meng: And that's the key insight that makes the whole thing work. Without that control, you're just guessing. With it, you have a baseline to measure against.
Tom: So we've got a paper that's asking a simple question with a deceptively simple title, but the implications are huge for privacy and accountability. Stick around, because we're going to dig into how they actually pull this off.
Jane: And trust me, the method is a lot more interesting than just flipping a switch. We'll get into that next.
Summary: Tom: So we've established that "SMA: Who Said That?" is asking where the model's answer actually comes from. Jane, can you break down what they actually did in the paper? How did they solve this?
Jane: So, they built a framework they call SMA, which stands for Source-aware Membership Audit. And the core trick is perturbation. They take your input, say a sentence, and they mess with it. They swap out keywords, they add weird unicode characters, they mask words out entirely.
Lu: And the idea is that these small changes should barely affect a model that's relying on its own internal knowledge. But they should drastically change the output if the model is relying on a specific retrieved passage. It's like testing how fragile the model's answer is.
Meng: Exactly. If you change one word in a question and the answer completely falls apart, that's a sign the model was leaning on a specific external source. If the answer stays solid, the model probably knew it all along.
Tom: So they're poking the model with a stick and seeing what makes it flinch. But how do they measure that flinch without seeing inside the model?
Jane: That's where the zero-gradient part comes in. They can't get gradients, so they use something called ridge regression. They create a bunch of these perturbed inputs, record the outputs, and then use math to estimate which input tokens were most responsible for the output.
Lu: It's a statistical approach. They're essentially building a map of cause and effect just by observing inputs and outputs. They don't need to know the model's architecture or its weights. They just need to see how the output changes when the input changes.
Meng: And they apply the same logic to images, which is the really wild part. For multimodal systems, they add Gaussian noise to an image, feed it to the model, and see how the description changes. That lets them attribute the output to the image or to the text prompt.
Tom: And then the final piece is the "RAG switch" we talked about. They run this whole attribution process with the retrieval system on, and then again with it off.
Jane: Right. And by comparing the attribution scores between those two runs, they can classify each piece of the output. If a word is important when retrieval is on but not when it's off, that word probably came from the retrieved document. That's a "Retrieved Member."
Lu: And if it's important in both cases, it's likely from the model's pre-training. That's a "Pretrained Member." And if it's not important in either case, it's probably just noise or novel user input. That's a "Non-Member."
Tom: So they're not just saying "this data was leaked." They're saying "this data was leaked from this specific pipe." That's a whole new level of detail.
Meng: And the numbers back it up. On text-based RAG systems, they improved accuracy by over fifteen percent compared to existing methods. And on the multimodal side, they were getting accuracy scores around seventy-nine percent, which is a huge jump from the baselines.
Jane: They even tested it on commercial models like ChatGPT and Gemini, which is impressive because those are truly black-box. You can't get any internal information from them at all.
Tom: So the summary is: they built a way to interrogate a model, figure out where its answers come from, and they proved it works across a bunch of different models and even with images. That's a solid contribution.
Lu: And it's a contribution that changes the conversation. We're moving from "did the model memorize this?" to "where did the model get this from?" That's a much more useful question for real-world auditing.
Tom: Alright, so we know what they did. But why does this matter for the real world? Let's talk about the improvements and implications next.
Improvements and Implications: Tom: So we've covered the method, the zero-gradient trick, the RAG switch. But what does this actually mean for people building and using these systems? Jane, what's the big takeaway for the industry?
Jane: The big takeaway is accountability. Right now, if a model spits out something sensitive, you have no idea where it came from. This paper gives you a tool to trace it back to the source. That's a game-changer for compliance.
Lu: And I'd push that even further, Tom. This isn't just about finding leaks after the fact. This enables proactive auditing. You can test your own system before you deploy it. You can check if your retrieval database is accidentally leaking private information, or if your model has memorized something it shouldn't have.
Meng: From an engineering standpoint, that's the dream. Imagine being able to run a test suite on your RAG pipeline that tells you exactly which documents are being over-relied upon. You could catch a bad document in the database before it causes a problem.
Tom: So it's a debugging tool as much as a security tool. That's a great point, Meng.
Jane: And there's a deeper implication here. The paper shows that existing membership inference attacks, the ones that just check if data was in training, they fall apart when RAG is introduced. The retrieval process destabilizes the input-output relationship they rely on.
Lu: Exactly. The old methods were built for a static model. But RAG makes the model dynamic. The same question can get different answers depending on what's in the database. SMA is the first method designed for that dynamic world.
Meng: And they showed that on commercial APIs, the baselines just fail. They need gradients or token probabilities that you can't get from ChatGPT. SMA only needs the text output, which makes it universally applicable.
Tom: So what are the downsides? There's got to be a catch.
Jane: The main catch is cost. This method requires a lot of queries. They're perturbing inputs dozens of times, sometimes up to a hundred times, to get good attribution. That's expensive if you're paying per API call.
Lu: And there's a sensitivity to hyperparameters. They found that the number of perturbations matters, and the noise level for images matters. Too much noise and the model gets confused. Too little and you don't see the effect.
Meng: But they also showed it's robust to the retrieval depth, the top-k setting. Even with more documents retrieved, their accuracy stayed high. That's a good sign for real-world use where you might have a large database.
Tom: So it's not a free lunch, but it's a tool that actually works. And that's more than most papers in this space can say.
Jane: And the authors are already thinking about the next steps. They mention using shadow models to run tests offline, which would cut the cost dramatically. That could make this a standard practice.
Lu: I think that's the real impact. This could become a standard part of the MLOps pipeline. You don't just test for accuracy and latency. You test for provenance. You test for source leakage.
Tom: So this paper is a stepping stone to a future where we hold these models accountable for what they say and where they got it from. That's a future I want to live in.
Jane: Me too, Tom. And speaking of the future, let's wrap this up and see what we're covering next.
Conclusion: Tom: Alright, we've spent a good chunk of time on "SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling," and I think it's fair to say this is one of those papers that reframes the problem.
Jane: Absolutely, Tom. We started with the simple question of "who said that?" and ended up with a framework that can answer it with surprising accuracy. They've shown that you can trace the source of generated content back to either the model's training data or an external retrieval database.
Lu: And they did it without needing any internal access. Just input-output pairs and some clever math. That's the part that impressed me the most. It makes the method practical for real-world systems.
Meng: Yeah, and the fact that they validated it on commercial models like ChatGPT and Gemini proves it's not just a lab experiment. It works on the stuff people actually use.
Tom: The implications are huge. For privacy, for compliance, for debugging. This isn't just an attack tool. It's an auditing tool that can make these systems safer.
Jane: And they've opened the door for future work. The cost is still high, but they've shown a path forward with shadow models. This is the kind of research that could become standard practice.
Tom: So, we're saying goodbye to SMA. It's been a fascinating look at how we can hold these black boxes accountable.
Jane: Definitely. And we're ready to move on to the next paper on the arXiv feed. We'll see what other surprises are waiting for us.
Tom: Thanks for listening, everyone. We'll catch you on the next one.
Shixuan Sun, Siyuan Liang, Jianjie Huang, Jingzhi Li, Xiaochun Cao
Sun Yat-sen University · Nanyang Technological University · University of Chinese Academy of Sciences · Zhongguancun Academy · Institute of Information Engineering, Chinese Academy of Sciences
cs.AI
Submitted: 2026-08-17
Updated: 2026-08-18
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 43/100
Key concepts
- Retrieval-Augmented Generation (RAG)
- A system where a model is enhanced by looking up fresh information from an external database or retrieval system. This allows the model to access knowledge beyond its original training data, which is the focus of the auditing problem.
- Source-aware Membership Audit (SMA)
- The core framework developed in the paper. It identifies whether a piece of output originated from the model's initial training set or if it was pulled live from an external retrieval database, providing specific accountability for content provenance.
- Perturbation and Attribution
- The method used to test the model. Researchers slightly alter input data (swapping keywords or adding noise) and observe how the output changes. This measures how fragile the answer is, indicating whether it relies on internal knowledge or a specific external source.
Terminology
Summary
Summary
This paper introduces the Source-aware Membership Audit (SMA) framework, the first membership inference method designed for fine-grained source attribution in Retrieval-Augmented Generation (RAG) and Multimodal Retrieval-Augmented Generation (MRAG) systems. The authors state: we propose the first Source-aware Membership Audit (SMA) that enables fine-grained source attribution of generated content in a semi-black-box setting with retrieval control capabilities.
The core problem addressed is that traditional Membership Inference Attacks (MIA) cannot reliably attribute generated outputs to their origins in RAG systems due to two main challenges: (1) Destabilized input-output associations
caused by dynamically incorporated retrieved content, and (2) Unobservable influence paths due to multimodal fusion
where images are encoded into latent representations. The paper notes: existing techniques struggle to accurately infer the origin of generated content.
The auditing task is formalized as a three-way classification problem, determining whether an input is a Pretrained Member (in the model's pretraining corpus), a Retrieved Member (in the external retrieval database), or a Non-Member (in neither). The authors state: This work shifts the focus of membership inference from 'whether the data has been memorized' to 'where the content is sourced from'.
The SMA framework operates in a semi-black-box environment with three assumed capabilities: Black-box access to the target system,
Retrieval control
(toggling RAG on/off), and Input perturbation.
The methodology consists of several key components:
-
Input Perturbation Design: For text, the method uses
keyword-level semantic perturbations
includingsynonym substitution, random masking, and word-level Unicode alteration.
For images, it injectspixel-wise Gaussian noise
with multiple standard deviation values (e.g., 10, 20, 30, 40). -
Zero-Gradient Auditing Mechanism: The paper proposes
a Zero-Gradient Attribution Mechanism to estimate the contribution of input tokens to the output under semi-black-box conditions.
This involves generating N randomly perturbed variants, associating each with a binary mask vector, and fitting a ridge regression model: β = (M⊤M + αΛ)−1M⊤r, where β represents token attribution scores. -
Attribution Scoring under RAG Switch: The method introduces the
Attribution Difference Score (ADS)
which measuresthe change in the impact of key words on the generated results by toggling the retrieval component.
The classification uses thresholds: words with β j ≥ τ are labeledPretrained Member,
while others are further classified using Diff(β j) with thresholds τ1 = 0.1 and τ2 = -0.1 to distinguishRetrieved Member
fromNon-Member.
-
Cross-Modal Attribution: For MRAG, SMA
projects image inputs into textual descriptions via MLLMs, enabling token-level attribution in the text modality, which for the first time facilitates membership inference on image retrieval traces in MRAG systems.
The experimental setup uses multiple datasets including ragbench and PubMedQA for rag storage
and WikiMIA and WikiMIA-24 datasets
for MIA comparison, plus VL-MIA-image and the Wikipedia image datasets for MRAG evaluation.
Baselines include LLaMA-2 7B, LLaMA-3.1 8B, Qwen2.5 7B, Qwen2.5-VL-7B, and commercial models ChatGPT-4o mini, ChatGPT-4.1 mini, and Gemini-2.5 flash.
Key results show SMA "outperforms state-of-the-art black-box MIA baselines in detecting source-specific membership leakage, with notable improvements in accuracy (+15.74%) and coverage metrics (+10.01%) under noise and zero-gradient conditions. In specific comparisons, SMA achieves
0.8624 accuracy and 0.5882 coverage on WikiMIA with LLaMA-2 7B, while baselines remain
below 0.53 on both metrics. For MRAG, SMA attains
the highest accuracy (0.7900) and AUC (0.8227) in Qwen2.5 VL 7B, while also achieving the lowest false positive rate (FPR) of 0.0785 among all methods."
Ablation studies demonstrate that adding noise consistently improves the accuracy for all models,
with improvements ranging from 0.0231 for LLaMA-3.1 8B to 0.1200 for Gemini-2.5 flash. The paper also notes that the gap between our semi-black-box method and the white-box upper bound is relatively narrow.
The paper discusses practical considerations including cost (The total token usage for each inference in SMA can be approximated as (TokenSMA = TokenOutput From Taret + 60)
), parameter sensitivity (optimal performance around 60–80 perturbations
), and limitations related to maximum token limit and sampling temperature
and high CPU and API usage.
The authors conclude: "we propose SMA (Source-aware Membership Auditing), the first membership auditing framework with source attribution capability, to determine whether the leaked content originates from the model's pre-training corpus or an external retrieved data source in a black-box setting."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:
Implementation: Add a three-way classification layer (Pretrained Member, Retrieved Member, Non-Member) to any RAG/MRAG system's output pipeline. This uses the paper's Attribution Difference Score (ADS) mechanism, which toggles the retrieval module on/off and compares token-level attribution via ridge regression.
Resulting capability: The AI system can now answer not just was this data in training?
but did this output come from the model's pre-training, from the external retrieval database, or from the user's prompt?
This enables content provenance tracing for compliance and privacy auditing.
Implementation: Integrate the paper's Zero-Gradient Auditing Mechanism (Section IV-C) as a middleware layer. This performs N=60–80 structured perturbations (Unicode character changes for text, Gaussian noise with σ∈ 10,20,30,40,50,60,80 for images), collects output responses, and fits a ridge regression model (β = (MTM + αΛ)−1MTr) to estimate token-level influence.
Implementation: Add a module that converts image inputs into textual descriptions via MLLM captioning, then applies the same token-level attribution in the text domain. This follows the paper's approach of using ỹ vis = f black(prompt, Ĩ) to project visual perturbations into the autoregressive text space.
Implementation: Implement the paper's finding that moderate Gaussian noise (σ≈50–60) improves attribution accuracy by 12–20% (as shown in Fig. 3b), while excessive noise (σ=80) degrades performance. Add an automatic noise-level selector that tests multiple σ values and picks the one maximizing attribution consistency.
Implementation: Add a control module that programmatically enables/disables the RAG/MRAG component and computes Diff(β j) = β(M)RAG(j) − β w/o(M)RAG(j) for each token. This directly measures each token's dependence on external retrieval versus internal knowledge.
Implementation: Adapt the SMA framework to work with API-only models (ChatGPT-4o mini, ChatGPT-4.1 mini, Gemini-2.5 flash) by using only textual responses and token-count-based response scoring (r(i) = γ1·Len(ỹ(i))/Len(x̃(i)) + γ2·Sim(ỹ(i), x̃(i))). This avoids the need for logits, gradients, or internal states that baselines (PETEL, Mink++, Min-K%PROB) require.
Implementation: Implement the paper's finding that 60–80 perturbations maximize accuracy with diminishing returns beyond that. Add a budget-aware scheduler that:
-
Starts with 60 perturbations
-
Monitors attribution stability (variance of β estimates)
-
Stops early if stable, or increases to 80 if unstable
-
Uses the paper's cost model (
Token SMA = Token Output + 60) to estimate API costs in advance
-
Audit any RAG/MRAG system (open-source or commercial API) to determine whether specific outputs leak data from pre-training, external retrieval, or user input—with 15.74% higher accuracy and 10.01% higher coverage than existing methods.
-
Trace image provenance in multimodal systems—a capability previously impossible—by attributing visual content to user-provided images, retrieved images, or model pre-training.
-
Provide per-token source labels (e.g.,
this word came from the retrieved document, this phrase came from pre-training
) enabling fine-grained content provenance reports. -
Detect retrieval database poisoning by identifying tokens whose attribution scores fluctuate abnormally when the retrieval module is toggled, flagging potentially malicious inserted content.
-
Perform continuous privacy compliance monitoring on deployed AI systems, even those behind commercial APIs, with predictable costs and without requiring internal model access.
-
Quantify retrieval dependence for each output, enabling data controllers to assess whether removing specific database entries would actually change system behavior—critical for GDPR
right to be forgotten
compliance.
Sources
- Captum: A unified and generic model interpretability library for PyTorch
- Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
- Using Captum to Explain Generative Language Models
- MM-LLMs: Recent Advances in MultiModal Large Language Models
- MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
- A Primer on Zeroth-Order Optimization in Signal Processing and Machine Learning
- Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
- Membership Inference Attacks against Large Vision-Language Models
- The Llama 3 Herd of Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen2.5 Technical Report
- Making Text Embedders Few-Shot Learners
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- PubMedQA: A Dataset for Biomedical Research Question Answering
- Qwen2.5-VL Technical Report
- LUMIA: Linear probing for Unimodal and MultiModal Membership Inference Attacks leveraging internal LLM states
- Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
- Feedback-Guided Extraction of Knowledge Base from Retrieval-Augmented LLM Applications
- Learning Transferable Visual Models From Natural Language Supervision
- "Why Should I Trust You?": Explaining the Predictions of Any Classifier
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection