Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs

summary

Video file (mp4)

The gist

This paper introduces VLM-UnBench, the first benchmark designed to rigorously evaluate training-free visual concept unlearning in Vision-Language Models (VLMs).

In short

The episode discusses a paper titled "Suppression Is Not Forgetting," which introduces VLM-UnBench, a benchmark for training-free visual concept unlearning in Vision-Language Models (VLMs). The hosts discuss how this benchmark rigorously tests whether suppressing concepts through prompts actually removes the underlying knowledge or just leads to compliance, concluding that verified methods are needed for safe deployment.

Key concepts

VLM-UnBench
This is the first benchmark designed to systematically test training-free visual concept unlearning in VLMs. It covers four levels of forgetting—object, scene, attribute, and privacy—across seven real-world datasets and eleven concept axes.
Residual Recoverability
The title suggests that even when using training-free suppression techniques, some residual knowledge remains that can be recovered later. This implies that complete erasure of a concept is not guaranteed by these methods alone.
Probe Taxonomy
This consists of three levels—direct identification, negation, and confirmation probes—used to test the model's response. It forces the model into different reasoning paths to determine if a concept is truly gone or just ignored under pressure.

Terminology used across episodes

This episode discusses

The paper

Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs · Read on arXiv

Zhangyun Tan, Zeliang Zhang, Susan Liang, Yolo Yunlong Tang, Lisha Chen Chenliang Xu

University of Rochester

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Suppression Is Not Forgetting".

Tom: This paper introduces VLM-UnBench, the first benchmark designed to rigorously evaluate training-free visual concept unlearning in Vision-Language Models (VLMs).

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, let's look at the title and who wrote this thing: "Suppression Is Not Forgetting: Residual Recoverability in Visual Concept Unlearning for VLMs." It’s quite direct about its main finding right off the bat.

Jane: And I think that title sets a very high bar because it suggests that even when we use these training-free suppression techniques, there's still some residual knowledge left behind that can be recovered later.

Lu: The authors are Zhangyun Tan, Zeliang Zhang, Susan Liang, Yolo Yunlong Tang, Lisha Chen, and Chenliang Xu from the University of Rochester. They come from a strong background in vision and language models development.

Meng: I'm curious how their work connects to the existing literature we've been looking at recently regarding methods like JUMP or those focusing on membership inference, since they are tackling something different—concept unlearning.

Lalam: It’s interesting that they focus on "Residual Recoverability"; that implies they aren't claiming perfect erasure, which is a very honest stance for a research paper in this field.

The paper's summary: Tom: They summarize the core problem as this gap: training-free unlearning methods suppress concepts through prompts without modifying model weights, but there’s no rigorous benchmark to tell us if that suppression is real knowledge removal or just temporary compliance.

Jane: Essentially, they’re asking if suppressing a concept through an instruction actually leads to genuine erasure of that visual information from the model's understanding.

Lu: To test this rigorously, they introduced VLM-UnBench, which covers four different levels of forgetting—object, scene, attribute, and privacy—across seven real-world datasets and eleven concept axes.

Meng: That structure sounds very comprehensive for testing across different types of visual knowledge retention issues. It moves beyond simple text checks into complex visual scenarios.

Lalam: I like that the benchmark is designed to test suppression across such diverse levels, from basic object labels to things like specific brand logos or personal identities, which is where the real practical problems lie for deployment.

The paper's improvements: Tom: Now let's talk about how they improve things; they propose VLM-UnBench as the first benchmark designed to systematically test this unlearning. They pair a three-level probe taxonomy with five specific evaluation conditions to separate true forgetting from just following an instruction.

Jane: That probe taxonomy, P1 through P3—direct identification, negation, and confirmation probes—is clever because it forces the model into different reasoning paths to see if the concept is truly gone or just ignored under pressure.

Lu: The five evaluation conditions they use are key: they include baseline normal settings, soft unlearning with realistic prompts, medium unlearning with stronger language, oracle hard settings where the target concept is explicitly revealed, and an oracle reverse setting using negation.

Meng: That setup seems designed to be very precise in separating compliance from erasure because it specifically includes conditions that reveal the target concept directly.

Lalam: This methodology is crucial because it directly addresses the paper's main concern: distinguishing between a model just learning to avoid an instruction and a model actually forgetting what the concept is.

Conclusion: Tom: So, wrapping up on this, the paper "Suppression Is Not Forgetting" really highlights that while training-free suppression is appealing for API deployments, we don't have solid proof it works as intended without a proper evaluation framework like VLM-UnBench.

Jane: It’s clear they show that current methods are often better at eliciting compliance when given oracle instructions than they are at actually removing the underlying visual concept knowledge.

Lu: The main implication here is that for anyone working on training-free unlearning, they need to move away from ad hoc testing and start using benchmarks that can isolate whether the model has genuinely forgotten something or just learned to avoid a specific prompt.

Meng: I think this means for us in engineering, we need to be much more careful when deploying these systems, especially if we are dealing with sensitive visual data where residual concepts could be problematic.

Lalam: For me, it reinforces the idea that if we want safe deployment of VLMs for things like proprietary brand monitoring or privacy protection, we can't rely on simple prompt suppression alone; we need a verified method of concept erasure.

More episodes

← Home