Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering".
Jane: The paper was written by Jieyuan Liu, Jianyang Gu, Shijie Chen, Jefferson Chen and Zhen Wang from University of California, San Diego and The Ohio State University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We're looking at a fascinating new paper called "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering."
Jane: That title sounds a bit ominous, Tom.
Tom: It really does, especially when you consider what "Primacy Bias" means for these AI systems.
Jane: If I'm understanding this correctly, the researchers from UCSD and Ohio State are saying the AI gets a bit too obsessed with whatever information it sees first.
Tom: Exactly, and they're looking at "Multimodal RAG," which is just a fancy way of saying the AI uses both images and text to find answers.
Jane: So, instead of just reading a textbook, it's like giving the AI a photo and a bunch of Wikipedia snippets to help it answer a question.
Tom: That's a great way to put it, Jane.
Lu: It's a massive problem because if the AI only listens to the first snippet, it might miss the perfect answer hidden later in the pile.
Tom: Lu, do you think this changes how we should be designing these visual assistants?
Lu: It definitely suggests we need to rethink how we present information to them so they don't just ignore the end of the prompt.
Meng: I wonder how this affects the actual reliability of the tools we're building for users.
Jane: That's the big question, Meng, because if the answer is in the tenth snippet but the AI only cares about the first one, the system fails.
Meng: It sounds like a huge headache for anyone trying to deploy this in a real-world app.
Lalam: It's a cultural shift too, because we're starting to realize that even when we give AI all the right facts, it might not actually "hear" them.
Tom: That's a deep thought, Lalam, and it leads us straight into what they actually found when they ran the experiments.
Paper discussion segment 2: Tom: We've been talking about that bias, and the actual data in "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering" is pretty startling.
Jane: They found that the performance doesn't follow that usual U-shape you see in text-only models.
Tom: Right, in text-only models, the AI usually remembers the start and the end, but forgets the middle.
Jane: But these multimodal models are different; they just seem to prefer the very beginning and then everything else drops off.
Tom: They saw accuracy plummet by sixteen to twenty-six points just because the correct answer was moved from the first slot to the last slot.
Jane: That's a massive gap, especially when you consider they tested models like the Qwen series and InternVL3.
Lu: It's wild because the researchers showed that adding an image actually makes this bias even stronger.
Tom: How much stronger are we talking about, Lu?
Lu: The paper says the multimodal setting amplifies that primacy bias by two to four times compared to just using text.
Meng: I'm looking at their "distractor-lift" error, and it seems like the AI is basically stealing answers from the wrong places.
Jane: Is that when it picks a wrong answer from the first passage instead of the right one at the end?
Meng: Yes, it's grabbing a piece of text from the very first snippet because it's so focused on that slot.
Lalam: It's like a student who only reads the first paragraph of a chapter and then tries to guess everything else.
Tom: That's a perfect analogy, and it brings us to the most important part: can we actually fix this?
Paper discussion segment 3: Tom: We're moving into the solutions part of "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering," and it's not as easy as it looks.
Jane: The researchers tried several common fixes on the retrieval side, like reranking the results.
Tom: They even tried something called MMR to diversify the snippets, but it didn't seem to help much.
Jane: It's because the reader models they used were "frozen," meaning they couldn't change how they process the information.
Tom: So, if the AI has a built-in habit of only looking at slot zero, reranking the snippets won't change that fundamental behavior.
Meng: That's frustrating because it means the people building the search engines are fighting a losing battle.
Jane: Exactly, Meng, because the problem isn't with the search; it's with how the AI reads what it finds.
Lu: I think the real path forward is training the readers themselves to be more attentive to every part of the prompt.
Tom: You're thinking about fine-tuning the models, right, Lu?
Lu: Precisely, we need to teach them that the gold answer could be anywhere in the list.
Meng: If we have to fine-tune every single model, that's going to be a massive engineering undertaking.
Lalam: But it's necessary if we want AI to be a truly reliable partner in how we explore knowledge.
Tom: It sounds like the researchers are calling for a shift from better searching to better reading.
Conclusion: Tom: We've covered a lot of ground today on "Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering."
Jane: It's a huge wake-up call for anyone building systems that combine vision and text.
Tom: We've learned that the primacy bias is real, it's amplified by images, and it's located right at that first prompt slot.
Jane: And most importantly, we've seen that fixing the search engine isn't enough; we have to fix the reader.
Lu: I'm honestly excited to see the next generation of models that are trained to handle these long, complex contexts without getting distracted.
Meng: I'll be watching to see if anyone releases a way to calibrate that attention without needing a massive retraining budget.
Lalam: I believe this will lead to a much more balanced way for AI to interact with the vast ocean of human information.
Tom: Well, that's all the time we have for this one.
Jane: Thanks for joining us, and we'll see you when we break down the next big paper.
Tom: Goodbye, everyone!
University of California, San Diego · The Ohio State University
cs.CL, cs.AI, cs.CV
Submitted: 2026-06-15
Updated: 2026-09-15
Comments: 20 pages, 8 figures. Accepted to EMNLP 2026 Main Conference; camera-ready version
Code: https://github.com/WeMWish/lost_at_the_end_code
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: This paper investigates position dependence in multimodal Knowledge-Based Visual Question Answering (KB-VQA) systems.
Key concepts
- Multimodal RAG
- A system where AI uses both images and text snippets to find answers to questions. Instead of relying solely on text, the model is provided with visual information and various data snippets to assist in its response process.
- Primacy Bias
- The tendency for AI models to focus disproportionately on information presented at the beginning of a prompt. In multimodal settings, this bias is amplified two to four times compared to text-only models, causing accuracy to drop if correct answers appear later.
- Distractor-lift error
- An error where an AI model selects an incorrect answer from the first snippet provided instead of finding the correct information located later in the sequence. This occurs because the model is overly focused on the initial information slot.
Terminology
Summary
This paper investigates position dependence in multimodal Knowledge-Based Visual Question Answering (KB-VQA) systems. While pure-text long-context LLMs exhibit a U-shaped lost-in-the-middle effect,
the authors examine whether this transfers to deployed multimodal RAG pipelines, where the ability to utilize retrieved context is critical for system reliability.
The Probing Protocol
The researchers designed the first controlled probe of reader-side position dependence in multimodal KB-VQA
using a gold-position protocol.
This method holds all variables bit-identical except for the index of the gold passage within a k-passage prompt, allowing for exact paired-bootstrap inference. The study was conducted at a scale representative of deployed pipelines, utilizing:
-
Three open-source 7B/8B VLM readers: Qwen2.5-VL-7B, InternVL3-8B, and Qwen3-VL-8B.
-
Two KB-VQA benchmarks: InfoSeek and E-VQA.
-
A multi-vector retriever (PreFLMR ViT-G) providing up to 50 candidate passages for k values up to 20.
The Primacy Effect
The study reveals that the position effect follows a monotonic primacy decay
rather than the U-shape observed in text-only long-context LLMs. The authors define primacy bias
as the reader’s systematic preference for evidence early in its prompt over otherwise-identical evidence at later positions.
Key findings include:
-
Gold-at-first beats gold-at-last by 16 to 26 points on every reader-by-benchmark cell.
-
The multimodal setting
amplifies an already-present text-mode primacy 2.2 to 4.5 times.
-
While the pattern is generally a decay, one exception (Qwen3-VL-8B × E-VQA) showed a
modality-induced sign reversal
where the last position slightly outperformed the middle.
The Locus of Error
To identify why readers fail at later positions, the authors performed targeted ablations to find the locus
of the effect. They observed that as gold position increases, failure modes shift from extraction-failed
to distractor-lift,
where the reader uses a wrong passage instead of the gold one. The investigation determined that:
-
The locus is
prompt slot 0 of the instruction-tuned reader.
-
Two ablations ruled out
image-token proximity and retrieval similarity as primary drivers.
-
A distractor-shuffle ablation proved the
reader follows the slot, not the similarity,
as the median PreFLMR rank of chosen distractors was 1, meaning they were often from slot 0 or 1.
Limitations of Retrieval-Side Fixes
The researchers tested several obvious retrieval-side responses
to see if they could recover performance, but found that they all leave the gap intact
on a frozen reader. The following interventions were evaluated:
-
MMR diversification.
-
Oracle reranking.
-
Rank-based distractor reordering.
Because these fixes provided no separable improvement,
the authors conclude that recall@k is the wrong metric for deployed KBVQA.
They argue that closing the gap requires reader-side intervention
rather than simply optimizing retrieval or increasing candidate pool sizes.
Improvements for AI systems
1. Position-Agnostic Multimodal Instruction Tuning
-
Improvement: Augment instruction-tuning datasets with multi-passage retrieval examples where the
gold
(correct) passage is systematically rotated through all available prompt slots (from index 0 to k-1) rather than being concentrated in a single top-1 or top- k format. -
Capability: The improved VLM will exhibit consistent retrieval accuracy regardless of where the relevant information appears in a long context, effectively neutralizing the
Lost at the End
primacy bias.
2. Distractor-Lift Penalty Objective
-
Improvement: Implement a specialized training loss that penalizes
distractor-lifting
—the specific error mode where the model extracts substrings from prompt slot 0 or 1 when the gold passage is located at a later index. -
Capability: The system will stop hallucinating answers based on highly similar but incorrect information found in early prompt positions, significantly reducing errors caused by topically near-gold distractors.
3. Reader-Aware Listwise Reranking
-
Improvement: Shift the reranker's optimization target from retrieval similarity (Recall@k) to
Reader Success
by training a reranker to specifically prioritize the gold passage into prompt slot 0 of the target VLM. -
Capability: The system will maximize its effective retrieval utility by aligning the retriever's output with the reader's positional attention bottleneck, ensuring that even if multiple candidates are retrieved, the most critical one is placed where the model is most likely to
see
it.
4. Inference-Time Attention Calibration
-
Improvement: Apply a dynamic attention re-weighting mechanism during inference that prevents excessive attention concentration on prompt slot 0 and redistributes the attention budget across all k passages.
-
Capability: This allows existing, frozen VLMs to utilize evidence found in middle and late context positions without requiring expensive retraining or fine-tuning.
Abstract
Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by conditioning a reader on passages retrieved from a Wikipedia-derived knowledge base. In pure-text long-context LLMs, retrieved-context use follows the U-shaped lost-in-the-middle effect of Liu et al. (2024): information at the start and end of context is used, the middle is lost. Whether this transfers to deployed multimodal KB-VQA is open. To close this gap, we design the first controlled probe of reader-side position dependence in multimodal KB-VQA: a gold-position protocol in which only the gold passage's prompt slot varies within question. We run it on three open-source 7B/8B VLM readers and two KB-VQA benchmarks with up to 20 retrieved passages. The shape flips from U to primacy: gold-at-first beats gold-at-last by 16 to 26 points on all six combinations of reader and benchmark, an effect we call Lost at the End; the gap holds at every scale we test, 3B to 32B, attenuating at 32B. Three targeted ablations narrow the cause. A text-only control that removes the image and changes nothing else shows the primacy is already present in text mode and does not depend on the image. Image-position and distractor-shuffle ablations trace the effect to prompt slot 0 of the instruction-tuned reader, where a second answer-bearing passage placed later is largely wasted. On a frozen reader, three retrieval-side fixes (MMR, oracle reranking, rank-based reordering) all fail to improve on the deployment default. Our findings indicate that recall@k is the wrong metric for deployed KB-VQA and that the remaining headroom sits on the reader side; we release our protocol as a controlled instrument for evaluating reader-side interventions.
Sources
- Qwen2.5-VL Technical Report
- mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
- Qwen3-VL Technical Report
- Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models
- Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering