Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
cs.CL, cs.CV, cs.IR
Submitted: 2025-10-28
Updated: 2026-09-16
Comments: EMNLP Main, Code here: https://github.com/alexmartin1722/mirage
Code: https://github.com/alexmartin1722/mirage
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: We introduce MiRAGE, an evaluation framework for retrieval-augmented generation (RAG) from multimodal sources.
Terminology
Abstract
We introduce MiRAGE, an evaluation framework for retrieval-augmented generation (RAG) from multimodal sources. As audiovisual media becomes a more prevalent source of information online, RAG systems must integrate such media into generation. Yet, existing evaluation methods for RAG are largely text-centric and do not readily transfer to multimodal settings. MiRAGE is a claim-centric approach to multimodal RAG evaluation, consisting of InfoF1, which assesses factuality and information coverage, and CiteF1, which assesses citation support and completeness. We show that, when applied by humans, MiRAGE strongly aligns with extrinsic judgments of output quality. We additionally introduce an automatic implementation of MiRAGE and compare it to multimodal variants of three prominent text-centric RAG metrics---ALCE, ARGUE, and RAGAS---finding that MiRAGE outperforms all three on text while being the only one to generalize to multimodal sources. We release open-source implementations and outline evaluation methods for multimodal RAG.
Sources
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
- Qwen2-Audio Technical Report
- Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
- Qwen2.5-VL Technical Report
- MegaWika: Millions of reports and their sources across 50 diverse languages
- CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
- MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation
- MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- Overview of the TREC 2023 NeuCLIR Track
- Overview of TREC 2024 Biomedical Generative Retrieval (BioGen) Track
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- VideoXum: Cross-modal Visual and Textural Summarization of Videos
- VideoRAG: Retrieval-Augmented Generation over Video Corpus
- Visual Instruction Tuning
- Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
- VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
- Core: Robust Factual Precision with Informative Sub-Claim Identification
- Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
- WikiVideo: Article Generation from Multiple Videos
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering