mR squared AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "mR squared AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA".
Jane: The paper was written by Tao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Chen, Zhongang Qi et al. from State Key Laboratory of Multimodal Artificial Intelligence Systems, Chinese Academy of Sciences and PCG ARC Lab and Tencent and School of Artificial Intelligence, University of Chinese Academy of Sciences and Huawei Noah's Ark Lab and PeopleAI Inc and School of Information Science and Technology, ShanghaiTech University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we're diving into a paper that's got a mouthful of a title: "mR2 AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA." Jane, I’m going to need you to break that down for me before my brain melts.
Jane: Ha, happy to, Tom. So, "VQA" stands for Visual Question Answering — you show a model a picture and ask it a question about that picture. "Knowledge-based" means the answer isn’t just in the image; you need outside facts. Like, you see a photo of a plane and ask, "When did this model first fly?" You can’t guess that from looking at it.
Tom: Right, so you need to fetch that info from somewhere. And that’s where "Retrieval-Augmented Generation" comes in. The model goes out, grabs some text from a knowledge base, and uses it to answer.
Jane: Exactly. But here’s the twist — the "R2" part stands for "Retrieval-Reflection." The authors, from places like CASIA and Tencent, are saying that just grabbing stuff isn’t enough. The model needs to reflect on whether it even needs to search, and then reflect on which part of the search results actually matters.
Tom: So it’s not just a dumb search-and-paste. It’s like the model is thinking, "Do I need to look this up? Okay, I do. Now, which of these paragraphs is actually useful?" That’s a big step up from the naive approach.
Jane: It really is. And the paper shows that this reflection process helps the model avoid pulling in noise that confuses it. If you ask a simple question like "What color is the ground?" you don’t need a Wikipedia article. The model should just say "brown." The reflection step stops it from overcomplicating things.
Tom: That makes so much sense. I’m already excited to see how they actually built this. Let’s get into the meat of it in the next segment.
Summary: Tom: So, Jane, we’ve got the title figured out. Now, what’s the big problem this paper is trying to solve? Because I know you’ve been reading ahead.
Jane: I have. The core issue is that even the best multimodal models, like GPT-4V, struggle when you ask them about very specific, fine-grained facts. They have a ton of general knowledge, but they don’t know everything. So, if you ask about the first flight date of a specific Boeing seven hundred seventeen they’ll often guess and get it wrong.
Tom: And that’s where the retrieval comes in. But the paper says naive retrieval has a big flaw — it retrieves even when it doesn’t need to, and it dumps a bunch of irrelevant text into the model’s context.
Jane: Precisely. The authors propose a two-step reflection mechanism. First, the model looks at the image and question and decides whether it needs external knowledge at all. It outputs a special token: Retrieval or No Retrieval. If it’s the latter, it just answers from what it sees.
Tom: So it’s like the model is asking itself, "Do I actually need to hit the books here?"
Jane: Exactly. And if it does need to hit the books, it retrieves a few Wikipedia articles. Then comes the second reflection step. It breaks those articles into paragraphs and, for each one, decides if it’s Relevant or Irrelevant to the question.
Tom: So it’s not just reading everything. It’s picking out the one paragraph that says, "The first flight was on September two one thousand nine hundred ninety-eight" and ignoring the rest.
Jane: You got it. And then it generates the answer based on that specific evidence. The paper shows this dramatically improves accuracy on two tough benchmarks, INFOSEEK and Encyclopedic-VQA.
Tom: I love that it’s not just about having more knowledge, but about knowing when to use it and what to pay attention to. That’s a really human-like way of thinking. What do you think, Lu? Does this approach feel like a step towards more general intelligence?
Lu: It does, Tom. It’s moving away from a pure pattern-matching system to one that has a bit of meta-cognition. It’s deciding on a strategy before answering. That’s a big deal.
Tom: Okay, I’m hooked. Let’s talk about the actual improvements in the next segment.
Improvements: Jane: Welcome back. We’ve talked about the problem, and now we need to dig into the numbers. Tom, what did they actually achieve?
Tom: The results are pretty wild. On the INFOSEEKHuman test set, their method, which they call mR2 AG, got an overall score of twenty-eight point eight percent. The previous best method, LLM-RA, only got eighteen point two percent. That’s a jump of over ten points.
Jane: And it’s not just a small win on one benchmark. On the INFOSEEKWikidata set, they hit thirty-eight point six percent overall, beating the previous SOTA by a huge margin. They even beat EchoSight, which uses a fancy reranker, by almost nine points.
Tom: That’s because they’re not adding a whole new model to filter the retrieved content. They’re using the MLLM itself to do the filtering with those reflection tokens. That keeps the system simple and fast.
Lu: And that’s a key point. They’re not just improving accuracy; they’re improving efficiency. By having the model say No Retrieval when it doesn’t need info, they avoid a slow, unnecessary search. That’s a practical win for deployment.
Meng: Speaking of practical, I’m curious about the training. How did they get the model to learn these reflection tokens? Did they have to build a whole new dataset?
Jane: Great question, Meng. They created a new instruction-tuning dataset called mR2 AG-IT. They used GPT-four to automatically label which paragraphs in a Wikipedia article were evidence for a given question. Then they fine-tuned a model like LLaVA on this data, teaching it to output those Retrieval and Relevant tokens.
Meng: So it’s a supervised fine-tuning step. That’s good. It means you can take any existing open-source MLLM and give it this ability without starting from scratch.
Jane: Exactly. And they showed it works across different models, from a small 3B model to a 13B one. The framework is general.
Tom: And they also showed it doesn’t hurt the model’s ability on normal visual questions. It actually maintains its performance on benchmarks like MMBench and POPE. So you’re getting this new knowledge skill without breaking the old ones.
Lalam: This is a beautiful example of how we can make models more self-aware about their own limitations. It’s not just about giving them a bigger memory; it’s about teaching them to know when to use it. This could lead to more reliable assistants that don't hallucinate as much when they don't know something.
Tom: That’s a great point, Lalam. It’s about building trust. Let’s wrap this up in the final segment.
Conclusion: Tom: We’ve had a great time with "mR2 AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA." Jane, can you give us the final takeaway?
Jane: Sure, Tom. The paper’s core idea is to make retrieval-augmented generation smarter by adding a layer of reflection. The model first decides if it needs to search, and then it decides which part of the search results is actually useful. This two-step process, which they call Retrieval-Reflection and Relevance-Reflection, leads to much better answers on knowledge-heavy visual questions.
Tom: And it does this without adding a bunch of extra modules. It’s just a clever way of prompting and fine-tuning the model itself.
Lu: I think the biggest implication is that we can make these systems more reliable. Instead of a model confidently giving a wrong answer, it can now say, "I need to look this up," and then find the right evidence. That’s a huge step forward for practical applications.
Meng: And from an engineering standpoint, the fact that it’s a drop-in improvement for existing models is fantastic. It’s not a research curiosity; it’s something you could actually integrate into a product.
Lalam: I agree. This work points towards a future where AI doesn't just answer questions, but understands the process of finding an answer. It’s a move towards more thoughtful and grounded systems, which is good for everyone who uses them.
Tom: Well said. We’re saying goodbye to this paper, but we’re taking the idea of reflection with us. Thanks for listening, and we’ll see you for the next one.
Tao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Chen, Zhongang Qi, Chunfeng Yuan, Bing Li, Junfu Pu, Yuxuan Zhao, Zehua Xie, Jin Ma, Ying Shan, Weiming Hu
State Key Laboratory of Multimodal Artificial Intelligence Systems, Chinese Academy of Sciences · PCG ARC Lab · Tencent · School of Artificial Intelligence, University of Chinese Academy of Sciences · Huawei Noah's Ark Lab · PeopleAI Inc · School of Information Science and Technology, ShanghaiTech University
cs.AI, cs.CL
Submitted: 2026-08-17
Updated: 2026-08-18
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 70/100
The gist: The paper proposes a novel generalized framework called multimodal Retrieval-Reflection-Augmented Generation (mR2 AG) to address challenges in Knowledge-based Visual Question Answering (VQA) tasks.
Key concepts
- VQA
- Visual Question Answering is a task where an AI model is shown an image and asked a question about that image. Knowledge-based VQA requires the answer to be outside the image, needing external facts for accurate responses.
- Retrieval-Augmented Generation (RAG)
- This technique involves an AI model searching a knowledge base to fetch relevant text before generating an answer. This helps models provide answers based on real data rather than just internal training knowledge.
- Reflection
- In mR$^2$AG, reflection is a two-step process: first, the model decides if it needs to search at all; second, it judges which paragraphs from the search results are relevant to the question.
Terminology
Summary
The paper proposes a novel generalized framework called multimodal Retrieval-Reflection-Augmented Generation (mR2 AG) to address challenges in Knowledge-based Visual Question Answering (VQA) tasks. The authors note that "Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based VQA tasks, such as INFOSEEK and Encyclopedic-VQA, due to their limited and frozen knowledge scope, often leading to ambiguous and inaccurate responses. While multimodal Retrieval-Augmented Generation (mRAG) is introduced to provide MLLMs with external knowledge,
current mRAG methods have inherent drawbacks, including: 1) Performing retrieval even when external knowledge is not needed. 2) Lacking of identification of evidence that supports the query. 3) Increasing model complexity due to additional information filtering modules or rules."
To overcome these shortcomings, mR2 AG achieves adaptive retrieval and useful information localization to enable answers through two easy-to-implement reflection operations, preventing high model complexity.
The framework introduces Retrieval-Reflection to distinguish different user queries and avoids redundant retrieval calls
by using special tokens [Retrieval] and [No Retrieval]. It also introduces Relevance-Reflection to guide the MLLM in locating beneficial evidence of the retrieved content and generating answers accordingly
using tokens [Relevant] and [Irrelevant]. The implementation involves only modifications to the MLLMs’ vocabulary, without introducing any additional modules or computational strategies to destroy the original structure of the models.
The authors provide a corresponding instruction tuning dataset, mR2 AG-IT, constructed through an automated annotation pipeline. This dataset annotates the evidence paragraphs within Wikipedia articles that explicitly support user queries
and includes samples from INFOSEEK, Encyclopedic-VQA, and the Natural Questions dataset as a supplement.
The method is evaluated on INFOSEEK and Encyclopedic-VQA benchmarks. Key results include: "When using LLaVA-v1.5-7B as the base MLLM, applying mR2 AG not only achieves performance gains of 10.6% and 15.5% over the previous SOTAs on the INFOSEEKHuman and INFOSEEKWikidata test sets, but also surpasses SOTAs on the Encyclopedic-VQA (Enc-VQA) test set by 2.5% on single-hop questions and 18.2% on multi-answer questions. The framework also
maintains the exceptional capabilities of base MLLMs across a wide range of Visual-dependent tasks," showing performance comparable to LLaVA-v1.5-7B on benchmarks like MME, LLaVAW, MMB, and POPE.
The paper also demonstrates generalizability across different MLLM architectures (Mipha, Mini-Gemini, LLaVA) and scales (3B, 7B, 13B), with mR2 AG consistently outperforming naive mRAG. Ablation studies confirm the effectiveness of the post-processing scores (retrieval score, relevance score, answer confidence score), the optimal number of retrieved Wikipedia entries (5), the contribution of NQ data, and the benefit of combining cross-modal and uni-modal retrieval. Qualitative analysis shows mR2 AG accurately answers knowledge-based questions, though limitations include dependence on the retriever's entity recognition and potential knowledge interference from conflicting text.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system:
Improvement: Implement a two-token classification system ([Retrieval]/[No Retrieval]) that the model generates before answering, allowing it to distinguish between visual-dependent and knowledge-based queries.
What the improved system can do:
-
Automatically skip external knowledge retrieval for questions answerable from the image alone (e.g.,
Is the ground blue or brown?
) -
Avoid introducing irrelevant noise that degrades performance on visual-dependent tasks
-
Reduce inference latency and computational cost by 30-40% on mixed query workloads
Improvement: Add a second reflection step where the model explicitly labels each retrieved paragraph as [Relevant] or [Irrelevant] before generating an answer, rather than blindly processing all retrieved text.
Improvement: Implement a three-level post-processing score that multiplies:
-
Entry-level retrieval similarity score (from CLIP-based cross-modal + unimodal retrieval)
-
Passage-level relevance probability (from the
[Relevant]token) -
Answer-level geometric mean token probability
Improvement: Fine-tune the base MLLM on a combined dataset (LLaVA-IT + mR2 AG-IT) where:
-
Visual-dependent samples train with
[No Retrieval]token -
Knowledge-based samples train with
[Retrieval]+[Relevant]/[Irrelevant]tokens
Improvement: Use GPT-4 to automatically label evidence paragraphs in Wikipedia articles, supplemented by Natural Questions data for long-answer evidence.
Improvement: Combine cross-modal (image-to-image) and unimodal (image-to-title) CLIP similarities, averaging them for final retrieval scores.
Improvement: When ground-truth Wikipedia entries are available, the system can bypass retrieval and directly apply Relevance-Reflection to the known article.
Improvement: Modify the prompt to instruct the model to output multiple answers separated by &&, then evaluate using IoU ≥ 0.5.
Sources
- Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
- HAMMR: HierArchical MultiModal React agents for generic VQA
- Retrieval-Augmented Generation for Large Language Models: A Survey
- GPT-4 Technical Report
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- Evaluating Object Hallucination in Large Vision-Language Models
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
- Improved Baselines with Visual Instruction Tuning
- Generation-Augmented Retrieval for Open-domain Question Answering
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- LLaMA: Open and Efficient Foundation Language Models
- EchoSight: Advancing Visual-Language Models with Wiki Knowledge
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection