Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

arXiv:2509.22415 · cs.CV, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models".

Jane: The paper was written by Jiawei Liang, Jianjie Huang, Xianghao Jiao, Siyuan Liang, Shiming Liu et al. from Sun Yat-sen University and Zhongguancun Academy and University of Chinese Academy of Sciences and Nanyang Technological University and Imperial College London.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the arXiv radio hour, everyone. I'm Tom, and today we're digging into a paper that's got a mouthful of a title — "Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models." Jane, I need you to translate that for our listeners, because I can barely say it without tripping.

Jane: Ha, I've got you, Tom. So the title is basically about figuring out *which part of an image* a multimodal AI model is actually looking at when it says a word. Like, if the model says "dog," we want to know if it's really looking at the dog, or if it's just guessing because the word "dog" is common. That's "visual attribution."

Tom: And the authors — Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, and Xiaochun Cao — they're coming from a bunch of places, including Sun Yat-sen University and the Chinese Academy of Sciences. That's a solid crew.

Jane: Yeah, and the key word in the title is "recomposition" and "residualization." I'll be honest, those sound like fancy math words, but the idea is pretty simple. The model processes the image in little chunks, like a grid of tiles. When it says "dog," it's reading those tiles. But the tiles are messy — they overlap, they mix together. So the paper says, let's look at the image multiple times with different tile arrangements, and then combine the results. That's the "recomposition" part.

Tom: And "residualization" is about cleaning up the noise. Because when the model says "dog," it's also thinking about the words that came before it, like "a big" or "the brown." Those previous words can pollute the map of where the dog is. So the paper subtracts that pollution out.

Jane: Exactly. It's like if you're trying to hear a guitar in a song, but the drums are too loud. You don't just turn up the guitar — you also turn down the drums. That's what this paper does for visual evidence.

Tom: And that's a big deal, because these multimodal models are getting used in everything from captioning photos to helping autonomous vehicles understand scenes. If we can't trust *why* they say something, we can't trust them in the real world.

Jane: Right. And the authors are claiming they can make that attribution much more reliable. We're going to get into the actual numbers and experiments in a bit, but I'm already excited because this feels like a real fix, not just a tweak.

Tom: Well, I'm hooked. Let's keep going and see what they actually did.

Summary of the Paper: Tom: So Jane, we've got the title down. Now let's talk about what the paper actually does. The abstract is dense, but the core problem is that when these multimodal models generate a word, the visual evidence — the part of the image they're "looking at" — is hard to inspect. There's a technique called "logit-lens" that tries to read the model's internal states and map them back to the image.

Jane: Right, and the problem is that this logit-lens reads each visual token — each little tile — independently. But the model doesn't process tiles independently. It mixes them together. So when you read them one by one, you get fragmented maps. The dog's face might be split across four tiles, and the map only lights up one of them.

Tom: And that's where "Evidence Recomposition" comes in. The authors take the same image, resize it or transform it in different ways, so the tiles land on different parts of the image. Then they read the attribution for each version and average them together. It's like taking multiple photos of the same scene from slightly different angles and combining them to get a clearer picture.

Jane: Exactly. And then there's the second problem — "Predictive Context Residualization." When the model says "dog," it's not just looking at the image. It's also processing the words that came before, like "a big." Those previous words have their own visual associations. So the attribution map for "dog" can get contaminated by the map for "big" or "a."

Tom: So they build a "context map" from all the preceding words, and then they subtract it from the current word's map. That's the residualization part. It's like removing the background noise from a recording.

Jane: And the results are pretty striking. On a model called Qwen2-VL-2B, they improve the F1-IoU score — that's a measure of how well the attribution map matches the actual object — from thirty-nine point one zero to forty-four point four five on the COCO Caption dataset. That's a big jump.

Tom: And they didn't just test one model. They tested LLaVA, Qwen2-VL, and InternVL, across different sizes, from 2B to 13B parameters. And the improvement holds everywhere. That's the kind of consistency you want in a method.

Jane: Yeah, and they also tested on three different datasets — COCO Caption, GranDf, and OpenPSG. So it's not just one lucky setup. The method generalizes.

Tom: I'm impressed. But I want to know more about how they actually implemented this. Is it expensive? Does it slow things down? Let's get into the improvements and the practical side.

Improvements Suggested by the Paper: Tom: Alright, Jane, so we've covered the what and the why. Now let's talk about the how — the actual improvements this paper suggests. And I want to bring in Lu and Meng for this, because they're going to have strong opinions.

Jane: Good idea. Lu, you're the researcher — what do you think makes this approach stand out?

Lu: Thanks, Tom. What I find most interesting is that the paper doesn't just add one trick. It identifies two distinct failure modes and then designs a specific fix for each. The Evidence Recomposition handles the grid problem, and the Predictive Context Residualization handles the context problem. That's a clean, principled approach. It's not just throwing a bigger model at the issue.

Meng: But Lu, I've got to ask about the cost. The paper mentions that Evidence Recomposition uses three different views of the image. That means three forward passes through the model, right? That's got to be expensive.

Jane: Actually, Meng, the paper addresses that. They report that the full method is only one point two zero times slower than the baseline TAM method. The extra views add some FLOPs — they go from three point five three to six point five one TFLOPs — but the memory footprint barely changes, from six point seven nine GB to six point nine six GB. So it's not a huge burden.

Meng: That's better than I expected. But I'm still worried about the practical deployment. If I'm running this on a real system, say for an autonomous vehicle, I can't afford a twenty percent slowdown on every single token.

Lu: That's a fair point, Meng. But the paper also shows that the ranking overhead — the context residualization part — is almost free. It's mostly the evidence views that cost. And they show that you can tune the number of views. The sensitivity analysis in Figure five shows diminishing returns after a moderate number of views. So you could probably get away with two views instead of three and still see most of the benefit.

Tom: And the improvements aren't just about accuracy on a benchmark. They also did perturbation tests — deletion and insertion. That's where you remove or add the highlighted regions and see if the model's confidence changes. ERCR performs better there too, which means the maps are actually capturing the evidence the model uses, not just matching human masks.

Jane: Right, and that's the key. A map can look pretty and match a human mask, but if the model doesn't actually rely on those regions, it's not a true explanation. The fact that ERCR passes the perturbation tests is a strong signal.

Meng: Okay, I'm getting more convinced. But I still want to know — does this work for all tokens, or just object words like "dog" and "car"? The paper talks about function words like "the" and "and" too.

Lu: Good question. The paper shows that PCR reduces the attribution mass more for function words than for object words. That's exactly what you want — function words shouldn't have strong visual evidence. So the method is doing the right thing across token types.

Tom: Alright, I think we've got a solid picture. Let's wrap this up and talk about what it all means.

Conclusion: Tom: So we've spent this whole episode on "Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models." Jane, give me your final take.

Jane: My final take is that this paper is a real step forward. It takes a known problem — that logit-lens attribution is messy and unreliable — and it fixes it with two clear, well-motivated ideas. Evidence Recomposition makes the maps less dependent on the grid, and Predictive Context Residualization cleans out the noise from previous words. The results are consistent across models and datasets.

Lu: And I'd add that the perturbation testing is what really sells it for me. The maps don't just look good; they actually reflect what the model is using. That's the difference between a pretty picture and a real explanation.

Meng: From an engineering standpoint, the one point two times slowdown is manageable, and the fact that you can tune the number of views gives you a nice accuracy-efficiency knob. I could see this being integrated into debugging tools for vision-language models.

Tom: And the broader impact — this matters for trust. If we're going to deploy these models in healthcare, in autonomous driving, in any high-stakes setting, we need to know *why* they say what they say. This paper gives us a better tool for that.

Jane: Absolutely. And the authors are already pointing toward future work — extending this to relation and action tokens, and maybe adapting the evidence views more intelligently. So this isn't the end of the story; it's a foundation.

Tom: Well said. That's it for this paper. We'll be back next time with something new. Thanks for listening, everyone.

Jane: See you soon.

Jiawei Liang, Jianjie Huang, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao

Sun Yat-sen University · Zhongguancun Academy · University of Chinese Academy of Sciences · Nanyang Technological University · Imperial College London

cs.CV, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 60/100

The gist: The paper proposes ERCR, an attribution framework for token-level visual evidence inspection in Multimodal Large Language Models (MLLMs), built from two components: Evidence Recomposition (ER) and

Key concepts

Visual Attribution
This refers to figuring out which specific part of an image a multimodal AI model is actually looking at when it generates a word. The goal is to determine if the model is correctly focusing on the relevant visual evidence or just guessing.
Evidence Recomposition
This technique addresses how models process images in chunks. It involves reading the image multiple times with different tile arrangements, then combining those results to create a clearer picture of where the model is looking.
Predictive Context Residualization
This method cleans up noise from preceding words. When a model generates a word, it can be influenced by previous words; this technique builds a context map from prior words and subtracts that influence to isolate the visual evidence for the current word.

Terminology

Summary

The paper proposes ERCR, an attribution framework for token-level visual evidence inspection in Multimodal Large Language Models (MLLMs), built from two components: Evidence Recomposition (ER) and Predictive Context Residualization (PCR).

The authors identify two practical instability sources in token-wise logit-lens attribution for MLLMs:

  1. Single-grid readout fragmentation: "Visual tokens are not independent local patches after MLLM processing; they contain context-mixed information from neighboring and global regions. Independently unembedding each token location can therefore assign broader regional evidence to a single visual tokenization grid, producing fragmented maps with limited object support."

  2. Preceding-token context interference: "The target token Tt is predicted from the image and preceding text tokens T<t. Under this autoregressive prediction context, the visual-token readout for kt can produce a map that shares spatial patterns with attribution maps of preceding tokens, including earlier words, punctuation, or connectors."

ER "constructs multiple evidence views with different token-to-region assignments and applies the same target-token readout to each view. It aggregates recurring target evidence to reduce dependence on a single visual readout grid. The recomposed attribution map is computed as a weighted least-squares consensus: ER is the weighted least-squares consensus map that favors evidence that recurs across views and reduces the influence of isolated view-specific peaks."

PCR uses RBO-based rank relevance to weight preceding-token attribution maps, forms a context map, and subtracts its fitted component from the ER map to mitigate preceding-token context interference. Specifically, PCR computes rank relevance between prediction distributions at preceding positions and the current target using Rank-Biased Overlap (RBO), converts lower rank relevance into normalized context weights, aggregates a preceding-token context map, fits a scalar coefficient via least-squares, and subtracts the fitted component from the ER map.

  1. Identification of two attribution-interface instabilities in logit-lens MLLM attribution: single-grid readout fragmentation from context-mixed visual tokens and preceding-token context interference.

  2. Proposal of Evidence Recomposition to reduce dependence on a single visual readout grid for the same generated target token.

  3. Introduction of Predictive Context Residualization, an RBO-based context residualization procedure that estimates preceding-token context and subtracts its fitted component from the ER map.

  4. Validation across LLaVA, Qwen2-VL, and InternVL families on three benchmarks, consistently outperforming existing attribution baselines.

  • On Qwen2-VL-2B, ERCR improves TAM F1-IoU from 39.10 to 44.45 on COCO Caption and from 30.83 to 37.20 on GranDf.

  • ERCR improves F1-IoU over TAM by 5.35, 6.37, and 3.69 points on COCO Caption, GranDf, and OpenPSG, respectively.

  • ERCR obtains the highest F1-IoU for every listed model and dataset across LLaVA1.5-7B/13B, Qwen2-VL-2B/7B, InternVL2.5-2B/4B/8B, InternVL3-2B, and InternVL3.5-2B.

  • Component ablation shows baseline without ER or PCR obtains 31.57 F1-IoU on Qwen2-VL-2B; ER only and PCR only improve it to 36.65 and 41.78 respectively, while their combination reaches 44.45.

  • In perturbation-based faithfulness evaluation, ERCR is strongest at the reported low budgets on all four evaluated models for deletion and insertion tests.

  • Efficiency analysis shows The full ERCR remains 1.20× over TAM in measured runtime with peak memory increasing from 6.79 GB to 6.96 GB.

The paper uses a TAM-compatible scoring protocol with three metrics: Obj-IoU (object words should localize corresponding image regions), Func-IoU (function words and low-visual-relevance tokens should avoid high-confidence image evidence), and F1-IoU as the harmonic mean, which penalizes methods that improve only one side of the evaluation. The paper also evaluates with fixed-prefix perturbation faithfulness using deletion and insertion tests.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems:

Implementation: Add an ERCR attribution layer to existing MLLM architectures (LLaVA, Qwen2-VL, InternVL) that:

  • Evidence Recomposition (ER): Runs the image through 3 scaled versions (0.5×, 0.75×, 1.0×), extracts visual-token hidden states from each, decodes target-token logits, and averages the resulting maps to reduce single-grid fragmentation.

  • Predictive Context Residualization (PCR): For each preceding text token, computes RBO-based rank relevance (top-50 vocabulary lists, decay p=0.8) against the target token, builds a weighted context map, fits a least-squares coefficient, and subtracts the fitted component before positive clipping and rank Gaussian filtering.

Resulting capability: The system can now produce token-level attribution maps that are 5–10 F1-IoU points more accurate than TAM across COCO Caption, GranDf, and OpenPSG, with better object localization (Obj-IoU) and lower false activation on function words (Func-IoU).

Sources

Related papers