PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement
cs.CR, cs.AI, cs.CV
Submitted: 2026-07-03
Updated: 2026-08-27
Comments: to appear in EMNLP 2026
Code: https://github.com/Zood123/PPE_Bench
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising privacy concerns.
Terminology
Abstract
Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising privacy concerns. Machine unlearning offers a way to remove such private knowledge without retraining from scratch. However, existing MLLM unlearning benchmarks have two major limitations. First, they rely on simplified images that contain only the single target individual, failing to reflect the visual complexity of real-world photos. Second, they typically assume that the forget set and retain set are fully separated, ignoring the fact that private information is often visually entangled with benign public information. For example, a private individual may appear with a public figure or in front of a well-known landmark, where unlearning the private target should not damage the public context. To address these limitations, we propose PPE-Bench, a new benchmark for evaluating MLLM unlearning under private-public entanglement. Each image contains a target individual to be forgotten and public information to be preserved, including public figure and landmark. We further introduce two simple but effective methods to better preserve public information during unlearning. Through experiments, we find that existing unlearning methods can reduce private information leakage, but often substantially harm adjacent public information.
Sources
- Qwen3-VL Technical Report
- The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
- GPT-4o System Card
- A Survey on Benchmarks of Multimodal Large Language Models
- Machine Unlearning in Generative AI: A Survey
- PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
- Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
- LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
- On Protecting the Data Privacy of Large Language Models (LLMs): A Survey
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs