Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs
Beining Xu, Hairui Wang, Jiaxin Wang, Changsheng Chen, Anirban Chakraborty
Shenzhen MSU-BIT University · Indian Institute of Science
cs.CV, cs.CR, cs.MM
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: ACM mm 2026
Code: https://github.com/xubeining/Beyond-Visual-Evidence
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper "Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs" investigates privacy risks in multimodal large language models (MLLMs) used for document
Terminology
Summary
The paper Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs
investigates privacy risks in multimodal large language models (MLLMs) used for document understanding, specifically in Key Information Extraction (KIE) tasks. The authors reveal that when input images lack sufficient visual evidence, these models often rely on memorized field relations from training data to infer missing content, thereby leaking multiple correlated fields containing sensitive personal information. This phenomenon is termed relational privacy leakage.
To address this, the paper makes three key contributions. First, it proposes the Dynamic Relational Unlearning Framework (DRUF), which comprises a Relational Decoupling Unlearning (RDU) module and a dynamic set update mechanism. DRUF suppresses the leakage of high-risk field pairs while preserving KIE performance. Second, it introduces DocPrivacyBench, a novel benchmark to systematically evaluate a model's susceptibility to privacy leakage under conditions of absent or minimal visual evidence. Third, it evaluates three MLLMs and six unlearning methods using this benchmark, assessing both post-unlearning leakage suppression and utility preservation.
The results demonstrate that existing MLLMs consistently exhibit privacy leakage when visual evidence is scarce, particularly on noisier datasets. In contrast, DRUF outperforms the strongest baseline by improving leakage suppression by 4.8 percentage points, effectively mitigating privacy risks while maintaining robust document information extraction performance.
The paper identifies that existing privacy protection methods are not well aligned with the document-image setting. Prior studies on multimodal privacy focus on general vision-language scenarios, analyzing leakage at the level of single attributes, individual outputs, or reconstructed content. These formulations overlook the structured nature of document-image KIE, where the main risk is whether multiple correlated sensitive fields are jointly leaked under weak visual evidence. Existing unlearning methods also face limitations: many suppress sensitive knowledge but do not explicitly preserve schema-level extraction ability, and they rely on predefined forget and retain sets, assuming forgetting targets are composed of individual concepts. In the paper's setting, leakage risk concerns private fields contained in samples across the entire training set, and leaked sensitive field pairs depend on current model behavior.
DRUF contains two key components. First, a dynamic forget-set update mechanism identifies leakage targets from the current model during probing and updates the forget set across multiple unlearning rounds, avoiding reliance on a static forget set. Second, in the Forget branch, Relational Decoupling Unlearning (RDU) shifts the forgetting target from isolated sensitive fields to exposed sensitive field pairs, directly penalizing the model's tendency to jointly generate correlated private fields.
The paper introduces DocPrivacyBench with two testing modes: Image Driven and Prompt Driven. In the Image Driven setting, images without valid visual evidence are used to evaluate privacy leakage while keeping the prompt fixed. In the Prompt Driven setting, 1,300 diverse prompts are automatically generated and divided into three categories: NIQP (Natural Interrogative Query Probing), SIEP (Structured Instructional Extraction Probing), and ECAP (Entity-Conditioned Associative Probing).
The experiments use three datasets: DocXPand-25k, IDNet, and IDNet(with noise). Three models are evaluated: LLaVA-1.5-hf, Xgen-Phi3, and Idefics2. Privacy leakage is measured using leakage accuracy (ACC) and leakage consistency (LC). The results show that LLaVA-1.5-hf exhibits markedly higher leakage risk than the other models. For example, on DocXPand-25k under Image Driven, its Acc@0.8 and Acc@1.0 reach 0.874 and 0.835, respectively, whereas Idefics2 achieves only 0.196 and 0.189. Across attack settings, Image Driven generally triggers more privacy leakage than Prompt Driven. From the perspective of training data conditions, lower data quality generally leads to higher privacy risk.
In the unlearning performance comparison, aggressive methods like GA and GA+KL reduce leakage rates to nearly zero but cause substantial drops in KIE capability. Methods like DPO, FTF, and DP preserve normal-task capability but still exhibit relatively high leakage rates. SCRUB achieves a balance but still yields a leakage accuracy of 0.049 under Image Driven. In contrast, DRUF reduces leakage rates under Prompt Driven and Image Driven to 0.000 and 0.001, respectively, while maintaining relatively high LC levels and KIE performance. The number of leaked pairs is also reduced from 62 in the base model to only 5 with DRUF.
The ablation study shows that privacy leakage persists under both irrelevant and training-similar visual conditions. Other Photos inputs lead to a higher leakage rate (0.955 at @1.0) but fewer unique leaked identities (32), while Face inputs lead to a lower leakage rate (0.825) but more unique identities (55). The supplementary material further analyzes prompt sensitivity, training process dynamics, test set size effects, and the sampling ratio γ, finding that γ=3 achieves the best trade-off between privacy protection and utility preservation.
Improvements for AI systems
Improvements to AI Systems:
-
Add a relational-leakage audit layer to document MLLMs. Before deployment, run the model on DocPrivacyBench-style inputs (images with no valid visual evidence, fixed prompts) and compute leakage accuracy (ACC) and leakage consistency (LC) for correlated field pairs (e.g., name + ID + address). If ACC > 0.1 or LC > 0.3, flag the model as high-risk and require retraining or unlearning.
-
Integrate Dynamic Relational Unlearning (DRUF) into the training pipeline. Replace static forget sets with a dynamic update mechanism that probes the current model each round to identify which sensitive field pairs are actually being leaked. This makes the unlearning process adaptive to model behavior rather than assuming fixed targets.
-
Shift forgetting targets from single attributes to correlated field pairs. Instead of unlearning
name
or "ID" in isolation, penalize the joint generation of any two or more private fields (e.g., name+address, ID+phone) using the Relational Decoupling Unlearning module. This directly reduces the risk of multi-field leakage, which is the primary threat in document KIE. -
Add a utility-preservation constraint during unlearning. Use a retain-set loss that explicitly enforces schema-level extraction ability (e.g., field names, layout structure) so that unlearning does not degrade KIE performance. This prevents the trade-off seen with aggressive methods like GA or GA+KL.
-
Implement a two-mode privacy evaluation harness for continuous monitoring. After any fine-tuning or unlearning update, run both Image Driven (no visual evidence) and Prompt Driven (diverse natural and structured queries) tests. Automatically generate 1,300 prompts across NIQP, SIEP, and ECAP categories to cover all plausible attack surfaces.
-
Add a data-quality-aware privacy risk estimator. Since lower training data quality (e.g., noisy labels) increases leakage risk, the system should estimate the noise level of the training set and adjust the unlearning intensity (e.g., sampling ratio γ) accordingly. For noisy datasets, increase γ to 3 for stronger suppression; for clean datasets, use lower γ to preserve utility.
-
Enable selective privacy hardening per input modality. The ablation shows that
Other Photos
inputs cause higher leakage rates but fewer unique identities, whileFace
inputs cause lower rates but more identities. The system should dynamically choose the unlearning target based on the input type—e.g., suppress high-rate pairs for photo-like inputs, and broaden the set of protected fields for face-like inputs. -
Add a post-unlearning verification step that checks both leakage and utility. After applying DRUF, the system should automatically verify that (a) leaked pair count drops below a threshold (e.g., ≤5 pairs, as achieved in the paper), and (b) KIE accuracy on a held-out clean set remains within 95% of the original model’s performance. If either fails, re-run unlearning with adjusted hyperparameters.
What the improved AI system can do:
-
Detect and prevent multi-field privacy leaks in document understanding tasks (e.g., invoices, IDs, medical forms) even when images are blurry, cropped, or missing key visual cues.
-
Maintain high KIE accuracy (e.g., field extraction, layout parsing) while reducing relational leakage from 62 leaked pairs to ≤5, and leakage accuracy from 0.87 to 0.001.
-
Adapt its privacy defenses dynamically based on the current model’s behavior, input image type, and training data quality—without requiring manual re-specification of forget sets.
-
Provide a standardized, reproducible privacy audit for any document MLLM, using both image-driven and prompt-driven attack simulations, so developers can certify compliance before deployment.
-
Balance privacy and utility automatically by tuning the unlearning sampling ratio (γ) and retain-set constraints, avoiding the catastrophic utility drops seen with aggressive unlearning methods.
-
Flag high-risk deployment scenarios (e.g., noisy training data, photo-like inputs) and trigger additional unlearning rounds or input filtering to reduce exposure.
Abstract
While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding MLLMs for identity document processing, this paper investigates the privacy issues inherent in Key Information Extraction (KIE) tasks. We reveal that when input images lack sufficient visual evidence, these models often rely on memorized field relations from training data to infer missing content, thereby leaking multiple correlated fields containing sensitive personal information. To mitigate this risk, we make three key contributions.First, we propose the Dynamic Relational Unlearning Framework (DRUF) which comprises a Relational Decoupling Unlearning (RDU) module and a dynamic set update mechanism. It suppresses the leakage of high-risk field pairs while preserving KIE performance.Second, we introduce DocPrivacyBench, a novel benchmark to systematically evaluate a model's susceptibility to privacy leakage under conditions of absent or minimal visual evidence.Third, we evaluate three MLLMs and six unlearning methods using this benchmark, assessing both post-unlearning leakage suppression and utility preservation.Our results demonstrate that existing MLLMs consistently exhibit privacy leakage when visual evidence is scarce, particularly on noisier datasets. In contrast, DRUF outperforms the strongest baseline by improving leakage suppression by 4.8 percentage points, effectively mitigating privacy risks while maintaining robust document information extraction performance.
Sources
- Unlearning Personal Data from a Single Image
- DocXPand-25k: a large and diverse benchmark dataset for identity documents analysis
- Extracting Training Data from Document-Based VQA Models
- PrIeD-KIE: Towards Privacy Preserved Document Key Information Extraction
- PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility
- A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
- Differentially Private Fine-tuning of Language Models
- Multi-PA: A Multi-perspective Benchmark on Privacy Assessment for Large Vision-Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models