Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models

summary

Video file (mp4)

The gist

This paper introduces PURGE (Partition-aware Unlearning for Removing spurious-correlation Generated Errors), a framework designed to mitigate spurious object-background correlations in Large

In short

The episode discusses a paper titled "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models." The hosts explain how the framework uses structured data partitioning to separate object-grounded examples from background correlations. They detail three partitioning methods and three unlearning objectives used to suppress background reliance while preserving object reasoning, concluding that this approach reduces hallucinations and improves model dependability.

Key concepts

Spurious Object-Background Correlations
This refers to when large vision-language models rely on irrelevant background cues instead of focusing on the actual object evidence. The paper addresses how these shortcuts lead to errors in the model's predictions.
Partition-Aware Unlearning
This is a framework designed to mitigate spurious correlations by separating data into retain and forget partitions. The goal is to suppress predictions driven by background cues while maintaining the model's ability to reason about objects.
Purge Methods (PURGE-D, PURGE-H, PURGE-M)
These are three complementary ways proposed to build the necessary data partitions. They include data-level disentanglement (PURGE-D), hybrid behavior-driven partitioning (PURGE-H), and model-informed semantic partitioning (PURGE-M).
Unlearning Objectives
These are mathematical tools used during parameter fine-tuning to remove unwanted knowledge. The authors use Gradient Ascent, KL-regularized unlearning, and Negative Preference Optimization to target specific dependencies in the model weights.

Terminology used across episodes

This episode discusses

The paper

Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models · Read on arXiv

Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan, Rhongho Jang, Dongxiao Zhu, Prashant Khanduri

Department of Computer Science, Wayne State University · Department of Electronics & Communications Engineering, Indian Institute of Information Technology Delhi (IIIT Delhi) · Institute for AI and Data Science at Wayne State University (AIDaS)

Large Vision-Language Models (LVLMs) achieve strong performance across many multimodal tasks; however, they often exploit spurious object-background correlations, resulting in predictions driven by contextual shortcuts rather than object-relevant visual evidence. Despite growing interest in hallucination and robustness evaluation, existing benchmarks provide limited control over whether model predictions are grounded in the target object or induced by correlated background cues. In this work, we introduce PURGE (artition-aware nlearning for emoving spurious-correlation enerated rrors), a framework for constructing, benchmarking, and mitigating spurious-correlation-induced failures in LVLMs. The framework consists of: -- (1) Structured dataset construction wherein we develop three complementary structured data construction strategies that partition examples by object-relevant evidence and spurious background cues, enabling controlled diagnosis of shortcut reliance; and -- (2) Partition-aware unlearning, which uses these partitions to selectively remove spurious object-background associations while preserving object-based reasoning. We evaluate the framework across multiple LVLMs, including LLaVA-1.6-7B, Qwen3-VL-8B-Instruct, and Qwen3.5-9B, together with CLIP as a vision-language encoder, on a diverse suite of benchmarks, including CHAIR, POPE, Causal-HalBench, MM-SpuBench, AMBER, MMHal, and Waterbirds. Our results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models".

Tom: This paper introduces PURGE (Partition-aware Unlearning for Removing spurious-correlation Generated Errors), a framework designed to mitigate spurious object-background correlations in Large Vision-Language Models (LVLMs).

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the specifics of this paper. The title is "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models," and it’s authored by Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan, Rhongho Jang, Dongxiao Zhu, and Prashant Khanduri.

Jane: Those authors are clearly experts in the multimodal space. The title points out that the main issue they're addressing is how Large Vision-Language Models often exploit spurious object-background correlations instead of looking at the actual object evidence.

Lu: What’s interesting about their approach, based on what I see here, is that they aren't just suggesting a fix for inference; they are proposing a framework for constructing, benchmarking, and mitigating these correlations right within the training or fine-tuning process.

Meng: So instead of just tweaking prompts during use to stop errors when they happen, this paper seems to suggest we can actually modify how the model learns those associations based on structured data separation. That implies a more fundamental change in the learning objective.

Lalam: That makes sense because if we can control *how* the model learns to separate what's object-related from what's just background noise, it should lead to much more stable and trustworthy outputs for our applications.

The paper's summary: Tom: Moving on, the core summary of "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models" is that they tackle hallucination by separating object-grounded examples from shortcut-prone ones using structured data partitioning.

Jane: Essentially, they argue that we need to create specific retain and forget partitions to get a controlled measure of how much the model is leaning on background cues versus actual object evidence. They propose three complementary ways to build these partitions: PURGE-D for data-level disentanglement, PURGE-H for hybrid behavior-driven partitioning, and PURGE-M for model-informed semantic partitioning.

Lu: That separation idea is key; by isolating the samples where contextual cues are present but the object is absent, they get a very specific set of examples to target during unlearning. It’s a very structured way to expose these spurious correlations that we usually just see as general errors.

Meng: So, it sounds like they aren't just looking at raw data; they are using model predictions and counterfactual inputs alongside ground-truth annotations to identify exactly where the model is failing due to those shortcuts. That’s a lot of engineering work.

Lalam: When you break it down, it means they are moving away from general debiasing methods that don't specify *why* a prediction failed, and instead focusing on isolating the specific association we want to remove.

The paper's improvements: Tom: Now for the improvements they suggest in "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models." They introduce partition-aware unlearning that aims to suppress predictions driven by background cues while keeping the object reasoning intact.

Jane: The mathematical goal they set is quite specific: they want the updated model parameters, theta', to satisfy P theta' (c z) to zero for the forget set, but still maintain P theta' (c o) about P theta (c o) for the retain samples. That’s a precise way of saying we want to purge the background reliance while keeping object-based reasoning stable.

Lu: The authors implement this using parameter-efficient fine-tuning with LoRA and three different unlearning objectives: Gradient Ascent, KL-regularized unlearning, and Negative Preference Optimization. This is a sophisticated mathematical tool applied directly to the model weights rather than just changing the input or output settings.

Meng: Applying that kind of targeted weight modification while keeping the overall model close to its original state on certain samples seems like a very clever way to achieve that balance between cleaning up errors and avoiding catastrophic forgetting on important object knowledge.

Lalam: The choice of those three unlearning objectives suggests they are hedging their bets, trying different mathematical levers to ensure they don't accidentally erase crucial object-based knowledge while successfully suppressing the unwanted background dependencies.

Conclusion: Tom: So we’ve covered the title, the summary of how it works, and those specific unlearning objectives. The overall conclusion of "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models" is that their framework consistently reduces hallucinations and spurious correlation errors across a diverse set of benchmarks while maintaining or improving overall performance in most settings.

Jane: Essentially, they show that by structuring the data separation and applying targeted unlearning, we can effectively reduce those background-driven predictions without significantly harming the model's ability to reason about objects. It’s a very practical result for making multimodal models more dependable.

Lu: The implication here is that we can start systematically diagnosing *where* a model is failing—is it missing the object, or is it just getting distracted by the scenery—and then surgically fix that specific dependency through partition-aware unlearning. That opens up new avenues for causal learning in vision and language.

Meng: From an engineering side, this means we can build iterative testing pipelines where we can specifically target these spurious correlations identified by the partitioning strategies, rather than relying on broad accuracy metrics alone to catch them later.

Lalam: I think the big picture is that this research gives us a pathway toward models that are not just smart pattern recognizers but are actually more robust reasoners grounded in verifiable visual evidence, which is exactly what we need for safe deployment.

More episodes

← Home