Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
summary
The gist
The paper proposes ERCR, an attribution framework for token-level visual evidence inspection in Multimodal Large Language Models (MLLMs), built from two components: Evidence Recomposition (ER) and
In short
The episode discusses a paper on improving visual attribution in multimodal large language models using 'Evidence Recomposition' and 'Predictive Context Residualization.' The hosts explain how these techniques fix issues with fragmented visual evidence and context noise, leading to significant accuracy improvements across various models and datasets. They conclude the method is a reliable step toward building trust in AI systems.
Key concepts
- Visual Attribution
- This refers to figuring out which specific part of an image a multimodal AI model is actually looking at when it generates a word. The goal is to determine if the model is correctly focusing on the relevant visual evidence or just guessing.
- Evidence Recomposition
- This technique addresses how models process images in chunks. It involves reading the image multiple times with different tile arrangements, then combining those results to create a clearer picture of where the model is looking.
- Predictive Context Residualization
- This method cleans up noise from preceding words. When a model generates a word, it can be influenced by previous words; this technique builds a context map from prior words and subtracts that influence to isolate the visual evidence for the current word.
Terminology used across episodes
This episode discusses
- Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models · Paper Radio
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Universal Camouflage Attack on Vision-Language Models for Autonomous Driving
- Efficient Backdoor Defense in Multimodal Contrastive Learning: A Token-Level Unlearning Method for Mitigating Threats
- From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
- Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Less is More: Fewer Interpretable Region via Submodular Subset Selection
- Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
- Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints · Paper Radio
- Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
- Can Attribution Predict Risk? From Multi-View Attribution to Planning Risk Signals in End-to-End Autonomous Driving · Paper Radio
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Analyzing Context Contributions in LLM-based Machine Translation
The paper
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models · Read on arXiv
Jiawei Liang, Jianjie Huang, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
Sun Yat-sen University · Zhongguancun Academy · University of Chinese Academy of Sciences · Nanyang Technological University · Imperial College London
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models".
Jane: The paper was written by Jiawei Liang, Jianjie Huang, Xianghao Jiao, Siyuan Liang, Shiming Liu et al. from Sun Yat-sen University and Zhongguancun Academy and University of Chinese Academy of Sciences and Nanyang Technological University and Imperial College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the arXiv radio hour, everyone. I'm Tom, and today we're digging into a paper that's got a mouthful of a title — "Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models." Jane, I need you to translate that for our listeners, because I can barely say it without tripping.
Jane: Ha, I've got you, Tom. So the title is basically about figuring out *which part of an image* a multimodal AI model is actually looking at when it says a word. Like, if the model says "dog," we want to know if it's really looking at the dog, or if it's just guessing because the word "dog" is common. That's "visual attribution."
Tom: And the authors — Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, and Xiaochun Cao — they're coming from a bunch of places, including Sun Yat-sen University and the Chinese Academy of Sciences. That's a solid crew.
Jane: Yeah, and the key word in the title is "recomposition" and "residualization." I'll be honest, those sound like fancy math words, but the idea is pretty simple. The model processes the image in little chunks, like a grid of tiles. When it says "dog," it's reading those tiles. But the tiles are messy — they overlap, they mix together. So the paper says, let's look at the image multiple times with different tile arrangements, and then combine the results. That's the "recomposition" part.
Tom: And "residualization" is about cleaning up the noise. Because when the model says "dog," it's also thinking about the words that came before it, like "a big" or "the brown." Those previous words can pollute the map of where the dog is. So the paper subtracts that pollution out.
Jane: Exactly. It's like if you're trying to hear a guitar in a song, but the drums are too loud. You don't just turn up the guitar — you also turn down the drums. That's what this paper does for visual evidence.
Tom: And that's a big deal, because these multimodal models are getting used in everything from captioning photos to helping autonomous vehicles understand scenes. If we can't trust *why* they say something, we can't trust them in the real world.
Jane: Right. And the authors are claiming they can make that attribution much more reliable. We're going to get into the actual numbers and experiments in a bit, but I'm already excited because this feels like a real fix, not just a tweak.
Tom: Well, I'm hooked. Let's keep going and see what they actually did.
Summary of the Paper: Tom: So Jane, we've got the title down. Now let's talk about what the paper actually does. The abstract is dense, but the core problem is that when these multimodal models generate a word, the visual evidence — the part of the image they're "looking at" — is hard to inspect. There's a technique called "logit-lens" that tries to read the model's internal states and map them back to the image.
Jane: Right, and the problem is that this logit-lens reads each visual token — each little tile — independently. But the model doesn't process tiles independently. It mixes them together. So when you read them one by one, you get fragmented maps. The dog's face might be split across four tiles, and the map only lights up one of them.
Tom: And that's where "Evidence Recomposition" comes in. The authors take the same image, resize it or transform it in different ways, so the tiles land on different parts of the image. Then they read the attribution for each version and average them together. It's like taking multiple photos of the same scene from slightly different angles and combining them to get a clearer picture.
Jane: Exactly. And then there's the second problem — "Predictive Context Residualization." When the model says "dog," it's not just looking at the image. It's also processing the words that came before, like "a big." Those previous words have their own visual associations. So the attribution map for "dog" can get contaminated by the map for "big" or "a."
Tom: So they build a "context map" from all the preceding words, and then they subtract it from the current word's map. That's the residualization part. It's like removing the background noise from a recording.
Jane: And the results are pretty striking. On a model called Qwen2-VL-2B, they improve the F1-IoU score — that's a measure of how well the attribution map matches the actual object — from thirty-nine point one zero to forty-four point four five on the COCO Caption dataset. That's a big jump.
Tom: And they didn't just test one model. They tested LLaVA, Qwen2-VL, and InternVL, across different sizes, from 2B to 13B parameters. And the improvement holds everywhere. That's the kind of consistency you want in a method.
Jane: Yeah, and they also tested on three different datasets — COCO Caption, GranDf, and OpenPSG. So it's not just one lucky setup. The method generalizes.
Tom: I'm impressed. But I want to know more about how they actually implemented this. Is it expensive? Does it slow things down? Let's get into the improvements and the practical side.
Improvements Suggested by the Paper: Tom: Alright, Jane, so we've covered the what and the why. Now let's talk about the how — the actual improvements this paper suggests. And I want to bring in Lu and Meng for this, because they're going to have strong opinions.
Jane: Good idea. Lu, you're the researcher — what do you think makes this approach stand out?
Lu: Thanks, Tom. What I find most interesting is that the paper doesn't just add one trick. It identifies two distinct failure modes and then designs a specific fix for each. The Evidence Recomposition handles the grid problem, and the Predictive Context Residualization handles the context problem. That's a clean, principled approach. It's not just throwing a bigger model at the issue.
Meng: But Lu, I've got to ask about the cost. The paper mentions that Evidence Recomposition uses three different views of the image. That means three forward passes through the model, right? That's got to be expensive.
Jane: Actually, Meng, the paper addresses that. They report that the full method is only one point two zero times slower than the baseline TAM method. The extra views add some FLOPs — they go from three point five three to six point five one TFLOPs — but the memory footprint barely changes, from six point seven nine GB to six point nine six GB. So it's not a huge burden.
Meng: That's better than I expected. But I'm still worried about the practical deployment. If I'm running this on a real system, say for an autonomous vehicle, I can't afford a twenty percent slowdown on every single token.
Lu: That's a fair point, Meng. But the paper also shows that the ranking overhead — the context residualization part — is almost free. It's mostly the evidence views that cost. And they show that you can tune the number of views. The sensitivity analysis in Figure five shows diminishing returns after a moderate number of views. So you could probably get away with two views instead of three and still see most of the benefit.
Tom: And the improvements aren't just about accuracy on a benchmark. They also did perturbation tests — deletion and insertion. That's where you remove or add the highlighted regions and see if the model's confidence changes. ERCR performs better there too, which means the maps are actually capturing the evidence the model uses, not just matching human masks.
Jane: Right, and that's the key. A map can look pretty and match a human mask, but if the model doesn't actually rely on those regions, it's not a true explanation. The fact that ERCR passes the perturbation tests is a strong signal.
Meng: Okay, I'm getting more convinced. But I still want to know — does this work for all tokens, or just object words like "dog" and "car"? The paper talks about function words like "the" and "and" too.
Lu: Good question. The paper shows that PCR reduces the attribution mass more for function words than for object words. That's exactly what you want — function words shouldn't have strong visual evidence. So the method is doing the right thing across token types.
Tom: Alright, I think we've got a solid picture. Let's wrap this up and talk about what it all means.
Conclusion: Tom: So we've spent this whole episode on "Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models." Jane, give me your final take.
Jane: My final take is that this paper is a real step forward. It takes a known problem — that logit-lens attribution is messy and unreliable — and it fixes it with two clear, well-motivated ideas. Evidence Recomposition makes the maps less dependent on the grid, and Predictive Context Residualization cleans out the noise from previous words. The results are consistent across models and datasets.
Lu: And I'd add that the perturbation testing is what really sells it for me. The maps don't just look good; they actually reflect what the model is using. That's the difference between a pretty picture and a real explanation.
Meng: From an engineering standpoint, the one point two times slowdown is manageable, and the fact that you can tune the number of views gives you a nice accuracy-efficiency knob. I could see this being integrated into debugging tools for vision-language models.
Tom: And the broader impact — this matters for trust. If we're going to deploy these models in healthcare, in autonomous driving, in any high-stakes setting, we need to know *why* they say what they say. This paper gives us a better tool for that.
Jane: Absolutely. And the authors are already pointing toward future work — extending this to relation and action tokens, and maybe adapting the evidence views more intelligently. So this isn't the end of the story; it's a foundation.
Tom: Well said. That's it for this paper. We'll be back next time with something new. Thanks for listening, everyone.
Jane: See you soon.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language