Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models
summary
The gist
This paper introduces PURGE (Partition-aware Unlearning for Removing spurious-correlation Generated Errors), a framework designed to mitigate spurious object-background correlations in Large
In short
The episode discusses a paper titled "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models." The hosts explain how the framework uses structured data partitioning to separate object-grounded examples from background correlations. They detail three partitioning methods and three unlearning objectives used to suppress background reliance while preserving object reasoning, concluding that this approach reduces hallucinations and improves model dependability.
Key concepts
- Spurious Object-Background Correlations
- This refers to when large vision-language models rely on irrelevant background cues instead of focusing on the actual object evidence. The paper addresses how these shortcuts lead to errors in the model's predictions.
- Partition-Aware Unlearning
- This is a framework designed to mitigate spurious correlations by separating data into retain and forget partitions. The goal is to suppress predictions driven by background cues while maintaining the model's ability to reason about objects.
- Purge Methods (PURGE-D, PURGE-H, PURGE-M)
- These are three complementary ways proposed to build the necessary data partitions. They include data-level disentanglement (PURGE-D), hybrid behavior-driven partitioning (PURGE-H), and model-informed semantic partitioning (PURGE-M).
- Unlearning Objectives
- These are mathematical tools used during parameter fine-tuning to remove unwanted knowledge. The authors use Gradient Ascent, KL-regularized unlearning, and Negative Preference Optimization to target specific dependencies in the model weights.
Terminology used across episodes
This episode discusses
- Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models · Paper Radio
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Qwen3-VL Technical Report
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
- Invariant Risk Minimization
- Learning Invariant Causal Mechanism from Vision-Language Models
- Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
- Debiasing Vision-Language Models via Biased Prompts
- Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
- Interpretable Debiasing of Vision-Language Models for Social Fairness
- A Closed-Form Solution for Debiasing Vision-Language Models with Utility Guarantees Across Modalities and Tasks
- Cross-Modal Attention Guided Unlearning in Vision-Language Models
- Decoupled Weight Decay Regularization
The paper
Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models · Read on arXiv
Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan, Rhongho Jang, Dongxiao Zhu, Prashant Khanduri
Department of Computer Science, Wayne State University · Department of Electronics & Communications Engineering, Indian Institute of Information Technology Delhi (IIIT Delhi) · Institute for AI and Data Science at Wayne State University (AIDaS)
Large Vision-Language Models (LVLMs) achieve strong performance across many multimodal tasks; however, they often exploit spurious object-background correlations, resulting in predictions driven by contextual shortcuts rather than object-relevant visual evidence. Despite growing interest in hallucination and robustness evaluation, existing benchmarks provide limited control over whether model predictions are grounded in the target object or induced by correlated background cues. In this work, we introduce PURGE (artition-aware nlearning for emoving spurious-correlation enerated rrors), a framework for constructing, benchmarking, and mitigating spurious-correlation-induced failures in LVLMs. The framework consists of: -- (1) Structured dataset construction wherein we develop three complementary structured data construction strategies that partition examples by object-relevant evidence and spurious background cues, enabling controlled diagnosis of shortcut reliance; and -- (2) Partition-aware unlearning, which uses these partitions to selectively remove spurious object-background associations while preserving object-based reasoning. We evaluate the framework across multiple LVLMs, including LLaVA-1.6-7B, Qwen3-VL-8B-Instruct, and Qwen3.5-9B, together with CLIP as a vision-language encoder, on a diverse suite of benchmarks, including CHAIR, POPE, Causal-HalBench, MM-SpuBench, AMBER, MMHal, and Waterbirds. Our results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models".
Tom: This paper introduces PURGE (Partition-aware Unlearning for Removing spurious-correlation Generated Errors), a framework designed to mitigate spurious object-background correlations in Large Vision-Language Models (LVLMs).
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about the specifics of this paper. The title is "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models," and it’s authored by Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan, Rhongho Jang, Dongxiao Zhu, and Prashant Khanduri.
Jane: Those authors are clearly experts in the multimodal space. The title points out that the main issue they're addressing is how Large Vision-Language Models often exploit spurious object-background correlations instead of looking at the actual object evidence.
Lu: What’s interesting about their approach, based on what I see here, is that they aren't just suggesting a fix for inference; they are proposing a framework for constructing, benchmarking, and mitigating these correlations right within the training or fine-tuning process.
Meng: So instead of just tweaking prompts during use to stop errors when they happen, this paper seems to suggest we can actually modify how the model learns those associations based on structured data separation. That implies a more fundamental change in the learning objective.
Lalam: That makes sense because if we can control *how* the model learns to separate what's object-related from what's just background noise, it should lead to much more stable and trustworthy outputs for our applications.
The paper's summary: Tom: Moving on, the core summary of "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models" is that they tackle hallucination by separating object-grounded examples from shortcut-prone ones using structured data partitioning.
Jane: Essentially, they argue that we need to create specific retain and forget partitions to get a controlled measure of how much the model is leaning on background cues versus actual object evidence. They propose three complementary ways to build these partitions: PURGE-D for data-level disentanglement, PURGE-H for hybrid behavior-driven partitioning, and PURGE-M for model-informed semantic partitioning.
Lu: That separation idea is key; by isolating the samples where contextual cues are present but the object is absent, they get a very specific set of examples to target during unlearning. It’s a very structured way to expose these spurious correlations that we usually just see as general errors.
Meng: So, it sounds like they aren't just looking at raw data; they are using model predictions and counterfactual inputs alongside ground-truth annotations to identify exactly where the model is failing due to those shortcuts. That’s a lot of engineering work.
Lalam: When you break it down, it means they are moving away from general debiasing methods that don't specify *why* a prediction failed, and instead focusing on isolating the specific association we want to remove.
The paper's improvements: Tom: Now for the improvements they suggest in "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models." They introduce partition-aware unlearning that aims to suppress predictions driven by background cues while keeping the object reasoning intact.
Jane: The mathematical goal they set is quite specific: they want the updated model parameters, theta', to satisfy P theta' (c z) to zero for the forget set, but still maintain P theta' (c o) about P theta (c o) for the retain samples. That’s a precise way of saying we want to purge the background reliance while keeping object-based reasoning stable.
Lu: The authors implement this using parameter-efficient fine-tuning with LoRA and three different unlearning objectives: Gradient Ascent, KL-regularized unlearning, and Negative Preference Optimization. This is a sophisticated mathematical tool applied directly to the model weights rather than just changing the input or output settings.
Meng: Applying that kind of targeted weight modification while keeping the overall model close to its original state on certain samples seems like a very clever way to achieve that balance between cleaning up errors and avoiding catastrophic forgetting on important object knowledge.
Lalam: The choice of those three unlearning objectives suggests they are hedging their bets, trying different mathematical levers to ensure they don't accidentally erase crucial object-based knowledge while successfully suppressing the unwanted background dependencies.
Conclusion: Tom: So we’ve covered the title, the summary of how it works, and those specific unlearning objectives. The overall conclusion of "Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models" is that their framework consistently reduces hallucinations and spurious correlation errors across a diverse set of benchmarks while maintaining or improving overall performance in most settings.
Jane: Essentially, they show that by structuring the data separation and applying targeted unlearning, we can effectively reduce those background-driven predictions without significantly harming the model's ability to reason about objects. It’s a very practical result for making multimodal models more dependable.
Lu: The implication here is that we can start systematically diagnosing *where* a model is failing—is it missing the object, or is it just getting distracted by the scenery—and then surgically fix that specific dependency through partition-aware unlearning. That opens up new avenues for causal learning in vision and language.
Meng: From an engineering side, this means we can build iterative testing pipelines where we can specifically target these spurious correlations identified by the partitioning strategies, rather than relying on broad accuracy metrics alone to catch them later.
Lalam: I think the big picture is that this research gives us a pathway toward models that are not just smart pattern recognizers but are actually more robust reasoners grounded in verifiable visual evidence, which is exactly what we need for safe deployment.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language