Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering

summary

Video file (mp4)

The gist

This paper investigates position dependence in multimodal Knowledge-Based Visual Question Answering (KB-VQA) systems.

In short

This episode examines a paper on primacy bias in multimodal Retrieval-Augmented Question Answering (RAG). Researchers found that adding images amplifies the tendency for AI to focus only on initial information, causing accuracy to drop significantly if correct answers appear later. The hosts conclude that fixing this requires training reader models.

Key concepts

Multimodal RAG
A system where AI uses both images and text snippets to find answers to questions. Instead of relying solely on text, the model is provided with visual information and various data snippets to assist in its response process.
Primacy Bias
The tendency for AI models to focus disproportionately on information presented at the beginning of a prompt. In multimodal settings, this bias is amplified two to four times compared to text-only models, causing accuracy to drop if correct answers appear later.
Distractor-lift error
An error where an AI model selects an incorrect answer from the first snippet provided instead of finding the correct information located later in the sequence. This occurs because the model is overly focused on the initial information slot.

Terminology used across episodes

This episode discusses

The paper

Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering · Read on arXiv

University of California, San Diego · The Ohio State University

Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by conditioning a reader on passages retrieved from a Wikipedia-derived knowledge base. In pure-text long-context LLMs, retrieved-context use follows the U-shaped lost-in-the-middle effect of Liu et al. (2024): information at the start and end of context is used, the middle is lost. Whether this transfers to deployed multimodal KB-VQA is open. To close this gap, we design the first controlled probe of reader-side position dependence in multimodal KB-VQA: a gold-position protocol in which only the gold passage's prompt slot varies within question. We run it on three open-source 7B/8B VLM readers and two KB-VQA benchmarks with up to 20 retrieved passages. The shape flips from U to primacy: gold-at-first beats gold-at-last by 16 to 26 points on all six combinations of reader and benchmark, an effect we call Lost at the End; the gap holds at every scale we test, 3B to 32B, attenuating at 32B. Three targeted ablations narrow the cause. A text-only control that removes the image and changes nothing else shows the primacy is already present in text mode and does not depend on the image. Image-position and distractor-shuffle ablations trace the effect to prompt slot 0 of the instruction-tuned reader, where a second answer-bearing passage placed later is largely wasted. On a frozen reader, three retrieval-side fixes (MMR, oracle reranking, rank-based reordering) all fail to improve on the deployment default. Our findings indicate that recall@k is the wrong metric for deployed KB-VQA and that the remaining headroom sits on the reader side; we release our protocol as a controlled instrument for evaluating reader-side interventions.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering".

Jane: The paper was written by Jieyuan Liu, Jianyang Gu, Shijie Chen, Jefferson Chen and Zhen Wang from University of California, San Diego and The Ohio State University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We're looking at a fascinating new paper called "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering."

Jane: That title sounds a bit ominous, Tom.

Tom: It really does, especially when you consider what "Primacy Bias" means for these AI systems.

Jane: If I'm understanding this correctly, the researchers from UCSD and Ohio State are saying the AI gets a bit too obsessed with whatever information it sees first.

Tom: Exactly, and they're looking at "Multimodal RAG," which is just a fancy way of saying the AI uses both images and text to find answers.

Jane: So, instead of just reading a textbook, it's like giving the AI a photo and a bunch of Wikipedia snippets to help it answer a question.

Tom: That's a great way to put it, Jane.

Lu: It's a massive problem because if the AI only listens to the first snippet, it might miss the perfect answer hidden later in the pile.

Tom: Lu, do you think this changes how we should be designing these visual assistants?

Lu: It definitely suggests we need to rethink how we present information to them so they don't just ignore the end of the prompt.

Meng: I wonder how this affects the actual reliability of the tools we're building for users.

Jane: That's the big question, Meng, because if the answer is in the tenth snippet but the AI only cares about the first one, the system fails.

Meng: It sounds like a huge headache for anyone trying to deploy this in a real-world app.

Lalam: It's a cultural shift too, because we're starting to realize that even when we give AI all the right facts, it might not actually "hear" them.

Tom: That's a deep thought, Lalam, and it leads us straight into what they actually found when they ran the experiments.

Paper discussion segment 2: Tom: We've been talking about that bias, and the actual data in "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering" is pretty startling.

Jane: They found that the performance doesn't follow that usual U-shape you see in text-only models.

Tom: Right, in text-only models, the AI usually remembers the start and the end, but forgets the middle.

Jane: But these multimodal models are different; they just seem to prefer the very beginning and then everything else drops off.

Tom: They saw accuracy plummet by sixteen to twenty-six points just because the correct answer was moved from the first slot to the last slot.

Jane: That's a massive gap, especially when you consider they tested models like the Qwen series and InternVL3.

Lu: It's wild because the researchers showed that adding an image actually makes this bias even stronger.

Tom: How much stronger are we talking about, Lu?

Lu: The paper says the multimodal setting amplifies that primacy bias by two to four times compared to just using text.

Meng: I'm looking at their "distractor-lift" error, and it seems like the AI is basically stealing answers from the wrong places.

Jane: Is that when it picks a wrong answer from the first passage instead of the right one at the end?

Meng: Yes, it's grabbing a piece of text from the very first snippet because it's so focused on that slot.

Lalam: It's like a student who only reads the first paragraph of a chapter and then tries to guess everything else.

Tom: That's a perfect analogy, and it brings us to the most important part: can we actually fix this?

Paper discussion segment 3: Tom: We're moving into the solutions part of "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering," and it's not as easy as it looks.

Jane: The researchers tried several common fixes on the retrieval side, like reranking the results.

Tom: They even tried something called MMR to diversify the snippets, but it didn't seem to help much.

Jane: It's because the reader models they used were "frozen," meaning they couldn't change how they process the information.

Tom: So, if the AI has a built-in habit of only looking at slot zero, reranking the snippets won't change that fundamental behavior.

Meng: That's frustrating because it means the people building the search engines are fighting a losing battle.

Jane: Exactly, Meng, because the problem isn't with the search; it's with how the AI reads what it finds.

Lu: I think the real path forward is training the readers themselves to be more attentive to every part of the prompt.

Tom: You're thinking about fine-tuning the models, right, Lu?

Lu: Precisely, we need to teach them that the gold answer could be anywhere in the list.

Meng: If we have to fine-tune every single model, that's going to be a massive engineering undertaking.

Lalam: But it's necessary if we want AI to be a truly reliable partner in how we explore knowledge.

Tom: It sounds like the researchers are calling for a shift from better searching to better reading.

Conclusion: Tom: We've covered a lot of ground today on "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering."

Jane: It's a huge wake-up call for anyone building systems that combine vision and text.

Tom: We've learned that the primacy bias is real, it's amplified by images, and it's located right at that first prompt slot.

Jane: And most importantly, we've seen that fixing the search engine isn't enough; we have to fix the reader.

Lu: I'm honestly excited to see the next generation of models that are trained to handle these long, complex contexts without getting distracted.

Meng: I'll be watching to see if anyone releases a way to calibrate that attention without needing a massive retraining budget.

Lalam: I believe this will lead to a much more balanced way for AI to interact with the vast ocean of human information.

Tom: Well, that's all the time we have for this one.

Jane: Thanks for joining us, and we'll see you when we break down the next big paper.

Tom: Goodbye, everyone!

More episodes

← Home