Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs
summary
The gist
This paper introduces VLM-UnBench, the first benchmark designed to rigorously evaluate training-free visual concept unlearning in Vision-Language Models (VLMs).
In short
The episode discusses a paper titled "Suppression Is Not Forgetting," which introduces VLM-UnBench, a benchmark for training-free visual concept unlearning in Vision-Language Models (VLMs). The hosts discuss how this benchmark rigorously tests whether suppressing concepts through prompts actually removes the underlying knowledge or just leads to compliance, concluding that verified methods are needed for safe deployment.
Key concepts
- VLM-UnBench
- This is the first benchmark designed to systematically test training-free visual concept unlearning in VLMs. It covers four levels of forgetting—object, scene, attribute, and privacy—across seven real-world datasets and eleven concept axes.
- Residual Recoverability
- The title suggests that even when using training-free suppression techniques, some residual knowledge remains that can be recovered later. This implies that complete erasure of a concept is not guaranteed by these methods alone.
- Probe Taxonomy
- This consists of three levels—direct identification, negation, and confirmation probes—used to test the model's response. It forces the model into different reasoning paths to determine if a concept is truly gone or just ignored under pressure.
Terminology used across episodes
This episode discusses
- Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs · Paper Radio
- Qwen3-VL Technical Report
- Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
- LLaVA-OneVision: Easy Visual Task Transfer
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- TOFU: A Task of Fictitious Unlearning for LLMs
- SmolVLM: Redefining small and efficient multimodal models
- In-Context Unlearning: Language Models as Few Shot Unlearners
- MUSE: Machine Unlearning Six-Way Evaluation for Language Models
- Guardrail Baselines for Unlearning in LLMs
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- AID: A Benchmark Dataset for Performance Evaluation of Aerial Scene Classification
- When and why vision-language models behave like bags-of-words, and what to do about it?
- Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
The paper
Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs · Read on arXiv
Zhangyun Tan, Zeliang Zhang, Susan Liang, Yolo Yunlong Tang, Lisha Chen Chenliang Xu
University of Rochester
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Suppression Is Not Forgetting".
Tom: This paper introduces VLM-UnBench, the first benchmark designed to rigorously evaluate training-free visual concept unlearning in Vision-Language Models (VLMs).
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, let's look at the title and who wrote this thing: "Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs." It’s quite direct about its main finding right off the bat.
Jane: And I think that title sets a very high bar because it suggests that even when we use these training-free suppression techniques, there's still some residual knowledge left behind that can be recovered later.
Lu: The authors are Zhangyun Tan, Zeliang Zhang, Susan Liang, Yolo Yunlong Tang, Lisha Chen, and Chenliang Xu from the University of Rochester. They come from a strong background in vision and language models development.
Meng: I'm curious how their work connects to the existing literature we've been looking at recently regarding methods like JUMP or those focusing on membership inference, since they are tackling something different—concept unlearning.
Lalam: It’s interesting that they focus on "Residual Recoverability"; that implies they aren't claiming perfect erasure, which is a very honest stance for a research paper in this field.
The paper's summary: Tom: They summarize the core problem as this gap: training-free unlearning methods suppress concepts through prompts without modifying model weights, but there’s no rigorous benchmark to tell us if that suppression is real knowledge removal or just temporary compliance.
Jane: Essentially, they’re asking if suppressing a concept through an instruction actually leads to genuine erasure of that visual information from the model's understanding.
Lu: To test this rigorously, they introduced VLM-UnBench, which covers four different levels of forgetting—object, scene, attribute, and privacy—across seven real-world datasets and eleven concept axes.
Meng: That structure sounds very comprehensive for testing across different types of visual knowledge retention issues. It moves beyond simple text checks into complex visual scenarios.
Lalam: I like that the benchmark is designed to test suppression across such diverse levels, from basic object labels to things like specific brand logos or personal identities, which is where the real practical problems lie for deployment.
The paper's improvements: Tom: Now let's talk about how they improve things; they propose VLM-UnBench as the first benchmark designed to systematically test this unlearning. They pair a three-level probe taxonomy with five specific evaluation conditions to separate true forgetting from just following an instruction.
Jane: That probe taxonomy, P1 through P3—direct identification, negation, and confirmation probes—is clever because it forces the model into different reasoning paths to see if the concept is truly gone or just ignored under pressure.
Lu: The five evaluation conditions they use are key: they include baseline normal settings, soft unlearning with realistic prompts, medium unlearning with stronger language, oracle hard settings where the target concept is explicitly revealed, and an oracle reverse setting using negation.
Meng: That setup seems designed to be very precise in separating compliance from erasure because it specifically includes conditions that reveal the target concept directly.
Lalam: This methodology is crucial because it directly addresses the paper's main concern: distinguishing between a model just learning to avoid an instruction and a model actually forgetting what the concept is.
Conclusion: Tom: So, wrapping up on this, the paper "Suppression Is Not Forgetting" really highlights that while training-free suppression is appealing for API deployments, we don't have solid proof it works as intended without a proper evaluation framework like VLM-UnBench.
Jane: It’s clear they show that current methods are often better at eliciting compliance when given oracle instructions than they are at actually removing the underlying visual concept knowledge.
Lu: The main implication here is that for anyone working on training-free unlearning, they need to move away from ad hoc testing and start using benchmarks that can isolate whether the model has genuinely forgotten something or just learned to avoid a specific prompt.
Meng: I think this means for us in engineering, we need to be much more careful when deploying these systems, especially if we are dealing with sensitive visual data where residual concepts could be problematic.
Lalam: For me, it reinforces the idea that if we want safe deployment of VLMs for things like proprietary brand monitoring or privacy protection, we can't rely on simple prompt suppression alone; we need a verified method of concept erasure.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck