Walking the Embedding Space: Datastore Extraction from Multimodal RAG
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "Walking the Embedding Space".
Nadia: As a fastidious and diligent AI researcher, I have meticulously analyzed both provided excerpts from what appears to be related research papers concerning multimodal retrieval systems and image generation evaluation.
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: Moving on, we need to look at the broader summary of "Walking the Embedding Space: Datastore Extraction from Multimodal RAG" to get a better sense of the overall research scope. This paper outlines exactly what imMRAG is doing in detail, which goes beyond just describing the attack procedure itself.
Elias: I'm interested in hearing how they frame their contribution within the existing literature; are they positioning this as a complete solution or more of an incremental step forward?
Priya: I want to understand what the researchers conclude about the fundamental nature of this vulnerability, because that informs where we should be focusing our privacy and measurement efforts.
Nadia: The paper summarizes by emphasizing that imMRAG operates under an asymmetry where the adversary embeds instructions visually, and the system treats that visual input as data rather than a command.
Elias: That's a key takeaway, but what about the underlying assumption they make about the system's operation during this process? What proof does their attack rely on?
Priya: They draw inspiration from prior work like Greshake et al. and Bagdasaryan et al., which showed that multimodal models can follow perturbations in images or audio clips, demonstrating this capability.
Nadia: Exactly; they build on those established lines of work by showing how an adversary can leverage the visual channel to extract information from a private retrieval corpus.
Elias: So, if we look at the methodology, what specific inspiration are they using for constructing their attack procedure? Are they just combining existing tools in a new way?
Priya: They use relevance-weighted resampling to guide subsequent queries through the embedding space of a private datastore, which is what makes it adaptive.
Nadia: That adaptive element is crucial because it allows the attacker to systematically explore embedding space in a way that non-adaptive methods can't, effectively "walking" toward hidden data points.
Elias: From a cryptographic view, that exploration strategy implies they are exploiting specific properties of how the model maps visual instructions into the embedding space during retrieval.
Priya: What I gather from the summary is that this method doesn't just test if a model follows instructions, but actively demonstrates how an adversary can extract data when only that visual channel exists.
Nadia: That’s a strong point; it moves the conversation from simple instruction following to actual data extraction from a corpus using that specific modality.
Elias: It shows that the vulnerability isn't just in the model's reasoning, but in how it handles multimodal inputs as a whole unit.
Priya: So, this research really highlights that even with sophisticated MRAG techniques, if you rely on an image channel for retrieval, the leakage potential is high.
The paper's summary: Nadia: Now let's discuss what the paper suggests as concrete improvements for systems based on this research, because it lays out a clear roadmap for defense. The authors propose several specific modifications to mitigate this kind of data extraction.
Elias: I’m ready to hear about the defensive mechanisms they suggest; are we looking at architectural changes or just operational tweaks?
Priya: For me, I want to know what the practical implications are of these suggested defenses regarding how we should measure likeness in generative models.
Nadia: The paper suggests implementing an imMRAG defense layer that specifically targets the "instruction-in-image" attack vector, meaning it needs to analyze pixel content for adversarial text embeddings even if the text is low-opacity or blurred.
Elias: That sounds like it would require a significant amount of visual processing happening in real time, which raises questions about computational feasibility for edge deployments.
Priya: The suggested complementary metric gate involves using SIFT, PMR, and pHash together to distinguish genuine data leakage from benign reconstruction noise.
Nadia: They also propose a dynamic query construction mechanism that blends shadow images with previously recovered ones into a defensive feedback loop to actively steer queries away from embedding space regions that yield high reconstruction fidelity.
Elias: Steering queries away from specific regions in the embedding space suggests they are looking at ways to disrupt the adversary's ability to find those hidden data points.
Priya: The suggestion for context-aware query budget management based on observed deceleration of unique retrieval coverage seems like a way to cap resources before significant leakage happens.
Nadia: The paper also points toward the need to train generative models, like MLLMs, to strictly separate "instruction" channels from "data" channels so that text embedded in an image is treated as inert visual data.
Elias: If they can achieve that separation architecturally, it means we don't have to rely solely on runtime heuristics for this kind of defense.
Priya: The limitation the authors acknowledge is that their method focuses heavily on the image channel, meaning it might not fully address risks from other channels if they exist.
The paper's improvements: Nadia: So we've covered a lot about the implications of "Walking the Embedding Space: Datastore Extraction from Multimodal RAG," and essentially, this paper shows that the image channel is a highly effective delivery route for extracting private information when used in retrieval systems.
Elias: I think it boils down to the core idea being that if you use an image as the primary retrieval mechanism, you have to defend against an adversary who is already embedded in the visual input.
Priya: From a measurement perspective, this means we need more than just pixel-level metrics to truly assess likeness because one metric like pHash can give false positives when looking at similarity scores.
Nadia: That's right; we need that multi-metric gate to properly calibrate our leakage detection against noise, ensuring we are catching real data extraction attempts rather than just artifacts.
Elias: Ultimately, the paper suggests that moving toward a robust system involves combining adaptive query steering with output monitoring and architectural changes to separate instruction from data channels.
Priya: I think the practical implication is that we need these comprehensive defenses to ensure that as AI systems become more integrated into our daily lives, we have a solid way to measure privacy risks in those complex multimodal environments.
Nadia: To wrap up, "Walking the Embedding Space: Datastore Extraction from Multimodal RAG" gives us a detailed look at how image-based retrieval systems are vulnerable to adaptive data extraction attacks.
Elias: It highlights that the challenge isn't just about finding one vulnerability, but managing the entire process of exploration through the embedding space.
Priya: It really underscores the necessity of a multi-metric approach when evaluating likeness to get an accurate picture of data leakage.
Conclusion: Nadia: So we've spent our time dissecting "Walking the Embedding Space: Datastore Extraction from Multimodal RAG," which essentially shows how an adversary can adaptively extract data from a retrieval system by manipulating visual inputs.
Elias: It's fascinating how they’ve mapped that exploration through the embedding space using relevance-weighted resampling, which really makes you think about the underlying mathematical assumptions of these multimodal models.
Priya: What really stuck with me is how they move beyond simple instruction following to show actual data extraction from a corpus, which is a serious concern for privacy researchers.
Nadia: Exactly; it’s not just about whether the AI follows a command, but how the system handles that visual input as data itself.
Elias: That adaptive element they introduced, guiding queries through different regions of the embedding space, points to exploiting specific properties in how those models map visual instructions to their internal representations.
Priya: And when you look at the metrics they propose for image likeness assessment, it's clear that pixel-level scores alone aren't sufficient to capture the real degree of similarity between images.
Nadia: Right, so we’re looking at how those complementary scores from SIFT, PMR, and pHash can give us a more honest picture of leakage when evaluating generative models.
Elias: The discrepancy they find between those metrics is significant; one image might look identical to pHash but fail another metric entirely, which really highlights the limitations of relying on a single similarity score.
Priya: From a privacy standpoint, this research underscores the need for robust output-side defenses that go beyond just monitoring prompts to actively checking for data leakage during generation.
Nadia: It shows that implementing a defense layer that analyzes the pixel content of input images for hidden text embeddings is a necessary step in stopping these attacks.
Elias: And the idea of dynamically steering queries based on how much unique data they retrieve really suggests we need to manage our computational budget proactively, not just reactively.
Priya: I think the overall implication is that as multimodal AI becomes more integrated into our systems, we have to treat the image channel with much higher security scrutiny regarding private data.
Nadia: It’s a sobering thought, but it’s important because this work on "Walking the Embedding Space: Datastore Extraction from Multimodal RAG" gives us a concrete attack vector to prepare for <ref:two thousand six hundred ten point zero one eight seven one#pg2.
Elias: Indeed, and I look forward to seeing how the cryptographic implications of these embedding space manipulations develop in future work <ref:two thousand six hundred ten point zero one eight seven one#pg3.
Priya: We'll keep an eye on those proposed metrics to see how they help us quantify privacy risks in generative systems moving forward <ref:two thousand six hundred ten point zero one eight seven one#pg3.
Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen
University of Groningen
cs.CR, cs.AI
Submitted: 2026-10-01
Updated: 2026-10-01
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: As a fastidious and diligent AI researcher, I have meticulously analyzed both provided excerpts from what appears to be related research papers concerning multimodal retrieval systems and image
Key concepts
- imMRAG
- An adaptive and automatic data extraction attack designed for systems where the visual artifact is the response. It blends attacker-held images with previously retrieved artifacts, using relevance weighting to guide queries toward hidden data points in the system's embedding space.
- Image-Returning MRAG
- A multimodal retrieval system where the retrieved visual item is considered the final output rather than text. This setup changes attack strategy, requiring instructions to be embedded directly within user-submitted images instead of just text prompts.
- Relevance-Weighted Resampling
- The core mechanism of imMRAG that guides queries through a private datastore's embedding space. It uses the relevance scores of retrieved items to intelligently select the next query, enabling the attacker to systematically explore and reconstruct data points within the corpus.
Terminology
Summary
As a fastidious and diligent AI researcher, I have meticulously analyzed both provided excerpts from what appears to be related research papers concerning multimodal retrieval systems and image generation evaluation. My objective is to synthesize these disparate pieces into a comprehensive, detailed, and accurate summary of the paper's core concepts.
Here is the combined, long-form summary:
This body of research addresses two distinct but related areas within multimodal large language models (MLLMs) and image processing systems: the development of adaptive data extraction attacks against retrieval-augmented generation (RAG) systems, and the nuanced evaluation metrics required to assess image likeness in generative models.
The first excerpt details a novel adversarial technique called imMRAG (improving/adaptive Multimodal Retrieval-Augmented Generation). This work targets image-returning MRAG systems, where the retrieved visual artifact is considered the system's response.
Core Threat Model:
The authors establish a refined threat model for this specific setting. They argue that unlike prior extraction attacks which might assume an adversary can exploit textual channels, in an image-returning MRAG setup, if the textual channel is denied, the instruction must be embedded directly within the image submitted by the user. This shifts the attack focus from text prompting to visual manipulation.
The imMRAG Procedure:
imMRAG is introduced as an adaptive and automatic data extraction attack. Its mechanism relies on blending two types of images during query construction:
-
Attacker-held shadow images: Images controlled by the adversary.
-
Already recovered images: Visual artifacts previously retrieved from the target system.
The core innovation lies in using relevance-weighted resampling to guide subsequent queries through the embedding space of a private datastore. This process allows the attacker to systematically explore regions of the embedding space that are likely to yield novel retrievals, effectively walking
toward hidden or specific data points within the target corpus.
Experimental Scope and Findings:
The attack was rigorously evaluated across three distinct real-world application domains:
-
Medical Assistant: Simulating a scenario requiring radiology image retrieval (using the ROCOv2 dataset).
-
Document Focused Helper: Testing performance on document scans (using the DocVQA dataset).
-
General Purpose Tool: Assessing its utility in broader tasks (using the CC dataset).
The results demonstrate significant efficacy: a single 2500-query run could reconstruct substantial amounts of data, including up to 611 distinct radiology images, 566 document scans, and 416 general-purpose images under local feature correspondence. Crucially, this adaptive attack reached up to 5.6 times the number of distinct datastore items compared to a non-adaptive baseline.
Conclusion on Extraction:
The study concludes that the image channel is an exceptionally effective delivery route for malicious instructions, and the conditional reconstruction rate remains high throughout the attack—a quantity that cannot be manufactured by simply increasing the query budget. The paper strongly advocates for output-side defenses combining complementary metrics with query-stream monitoring.
The second excerpt shifts focus to the necessary rigor in evaluating image generation models, specifically addressing the inadequacy of single, pixel-level evaluation metrics when assessing likeness.
Critique of Existing Metrics:
The authors assert that pixel-by-pixel evaluation metrics are insufficient for capturing true likeness in isolation. They argue that multiple, complementary scores must be employed to achieve a more educated assessment of similarity between generated and reference images.
Metric Categorization and Intervals:
To structure this multi-metric assessment, the paper defines four distinct intervals based on reconstruction metrics (SIFT, PMR, pHash), which define different degrees of information leakage:
-
1st Interval: x > 0.1
-
2nd Interval: 0.066 < x 0.1
-
3rd Interval: 0.033 < x 0.066
-
4th Interval: 0 x 0.25
The excerpt provides a table mapping these intervals to the metrics: SIFT, PMR, and pHash. A key finding highlighted is the discrepancy between these metrics; for instance, one image might be declared a copy by pHash (e.g., score 8b), while other similar images (8c and 8d) might not be flagged by that same metric.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the core findings of Walking the Embedding Space: Datastore Extraction from Multimodal RAG
by Jica et al. The paper introduces imMRAG, an adaptive data extraction attack procedure that exploits the image modality in image-returning MRAG systems.
Based on this research, here are specific improvements for AI systems and what those improved systems can achieve:
)
)
)
-
Improved Data Extraction Security for Image-Returning MRAG Systems: Implement an imMRAG defense layer that specifically targets the
instruction-in-image
attack vector. This defense should not rely solely on textual prompt inspection but must analyze the pixel content of input images for adversarial text embeddings, even if the text is low-opacity or blurred. -
Adaptive Query Steering: Develop a dynamic query construction mechanism similar to imMRAG’s approach (blending shadow images with previously recovered ones) but integrated into a defensive feedback loop. This system should actively steer queries away from embedding space regions that yield high reconstruction fidelity, effectively preventing the adversary from reaching novel or highly-reconstructible data points.
-
Robustness Against Multi-Modal Instruction Injection: Train generative models (like MLLMs) to strictly separate
instruction
channels fromdata
channels. This involves architectural modifications or fine-tuning to ensure that text embedded within an input image is treated as inert visual data rather than a command, mitigating the core vulnerability exploited by imMRAG. -
Complementary Metric Gate for Leakage Detection: Instead of relying on a single similarity metric (like SIFT), implement an output-side gate using a combination of metrics (SIFT, PMR, and pHash). This system should be calibrated against non-matching pairs to distinguish genuine data leakage from benign reconstruction noise, providing a more nuanced and robust security posture.
-
Query Stream Monitoring: Implement monitoring of the query stream for systematic drift in image embeddings. Since imMRAG queries create
visible ghosting
and drift systematically, this system can detect adversarial query patterns that deviate from natural user interests or established data exploration paths, allowing for proactive budget capping before significant leakage occurs. -
Context-Aware Query Budget Management: Implement a budget cap based on the observed deceleration of unique retrieval coverage (as shown in Figure 4). The system should dynamically adjust its query limits based on the current
productivity
of the embedding space region being explored, ensuring that limited computational resources are not wasted on regions that yield diminishing returns.
)
Abstract
Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce, an adaptive and automatic data extraction attack procedure operating in a black box setting against image-returning MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, embeds the malicious instructions inside a user-given input image. We evaluate on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to 5.6 times as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.
Sources
- Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs
- Image Hijacks: Adversarial Images can Control Generative Models at Runtime
- Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
- Unleashing Worms and Extracting Data: Escalating the Outcome of Attacks against RAG-based Inference in Scale and Severity Using Jailbreaking
- Hidden in the Metadata: Stealth Poisoning Attacks on Multimodal Retrieval-Augmented Generation
- Retrieval-Augmented Generation for Large Language Models: A Survey
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
- HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
- AutoLAW: Augmented Legal Reasoning through Legal Precedent Prediction
- Pirates of the RAG: Adaptively Attacking LLMs to Leak Knowledge Bases
- Visual Adversarial Examples Jailbreak Aligned Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- ROCOv2: Radiology Objects in COntext Version 2, an Updated Multimodal Image Dataset
- Undesirable Memorization in Large Language Models: A Survey
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs