MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models".
Tom: The gist: MMLongCite introduces a comprehensive benchmark designed to evaluate the fidelity of large vision-language models in long-context scenarios by requiring citation generation across diverse multimodal contexts.
Jane: First, who's behind it and why it matters.
Paper summary: Jane: So, looking at the whole picture of MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models, what’s the big picture implication for us in the AI space?
Tom: Well, it points out that simply making context windows bigger isn't enough; we need to focus on improving how models actually utilize that long context reliably across different senses and formats.
Lu: This benchmark provides a much more rigorous standard than what’s currently available, moving the evaluation from just text faithfulness to multimodal faithfulness in long contexts. It shows where the current state-of-the-art LVLMs fall when faced with this kind of challenge, as detailed in MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models.
Meng: For practical impact, it gives us concrete metrics—like Citation Recall and Precision—to see not just if the answer is right, but how well it actually backs up that answer with the provided long context. That’s a huge step for building trustworthy applications.
Jane: And what they do by including tasks like MMLongCite-Grounding, which tests visual grounding at different difficulty levels, is show how models struggle when the visual information gets dense or complex in those long sequences.
Tom: It highlights that even with long context, dense visual information can cause a decline in both grounding accuracy and answer correctness, which is a sobering finding from MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models.
Lu: And the authors also pointed out some specific issues they found, like the "lost-in-the-middle" problem for certain models when dealing with images in the center of a very long context. It gives us specific areas to focus our research on, as MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models.
Lalam: From a cultural standpoint, if we can build systems that are faithful to long multimodal contexts, it means the AI we deploy will be able to handle much more complex, real-world information streams without making up its sources. That’s a big deal for how people interact with these new tools.
Meng: So, the overall message of MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models is that extending context length needs to be paired with better mechanisms for reliable information use and attribution, not just brute force input size.
Tom: Exactly. It’s a call to action for research to focus on the core mechanisms of how these models process and attribute information within those massive multimodal inputs.
Conclusion: Tom: So, MMLongCite is basically a new test for big vision language models to see if they actually stick to what they see when you give them really long documents or videos.
Jane: That’s right, Tom. The authors are Keyan Zhou and his team, and they've built this whole benchmark structure to check that fidelity in these long-context situations.
Lu: What's interesting is how they structured it—they aren't just looking at one thing; it covers different types of visual inputs like images, videos, and text mixed together.
Meng: I’m curious about the practical side. The results show a real gap between an AI giving you a right answer and actually citing the right source in that long context.
Lalam: That decoupling is what they highlight most—models can get the facts correct but fail at providing accurate attribution for those facts within a massive input.
Tom: Exactly, Lalam. It shows that just having a long memory isn't enough; you need the mechanism to reliably pull information from that memory correctly.
Jane: And looking at the final results, they found that context length alone doesn't solve everything; where the important stuff is actually located in that long sequence matters a lot.
Lu: The analysis pointed out a "lost-in-the-middle" problem for some models when the crucial visual information was buried deep inside a very long context window.
Meng: That makes sense from an engineering standpoint; if the model’s attention mechanism gets overloaded by everything in between, it starts missing the specific detail it needs.
Lalam: It suggests that future work needs to focus less on just stretching the input and more on improving how these models prioritize and retrieve relevant visual evidence within a long stream.
Tom: So, MMLongCite isn't just another dataset; it’s setting a new standard for how we judge if these large vision language models can actually be trusted with long-form information.
Jane: It really frames the challenge as moving past simple input size and focusing on smarter information utilization within that size.
Lu: And that leads us right into what the authors suggest next—how they plan to extend this benchmark to cover even more data types beyond just vision.
Keyan Zhou, Zecheng Tang, Lingfeng Ming, Guanghao Zhou, Qiguang Chen, Dan Qiao, Zheming Yang, Libo Qin, Minghui Qiu
Soochow University 2ByteDance Department of Computer Science Harbin Institute of Technology Central South University
cs.CV, cs.CL
Submitted: 2025-10-15
Updated: 2026-10-05
Code: https://github.com/jiqimaoke/MMLongCite
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: The gist: MMLongCite introduces a comprehensive benchmark designed to evaluate the fidelity of large vision-language models in long-context scenarios by requiring citation generation across diverse
Key concepts
- MMLongCite Benchmark
- This is a comprehensive test designed to evaluate the 'fidelity' of LVLMs in long-context situations. It includes eight distinct tasks covering six different context length intervals and incorporates various data types (text, images, videos) to rigorously check model performance.
- Context Modalities
- The benchmark tests models using three main context formats: image-only, interleaved image-text (mixing both), and video-only. This diversity ensures the evaluation covers how well models process different kinds of long multimodal inputs, moving beyond text-only testing.
- MMLongCite-Grounding
- This specific task measures a model's ability to locate objects accurately within visual contexts. It has two settings: 'Easy,' where models see composite images made from four smaller sources, and 'Hard,' where they must process one single, very large composite image.
- Citation Fidelity Metrics
- These metrics (Recall, Precision, and F1) measure the quality of the information a model cites. The study found that models often give correct answers but fail to provide accurate sources, showing a disconnect between getting the right answer and properly attributing it.
Terminology
Summary
The gist: MMLongCite introduces a comprehensive benchmark designed to evaluate the fidelity of large vision-language models in long-context scenarios by requiring citation generation across diverse multimodal contexts.
MMLongCite Benchmark Design
MMLongCite is a comprehensive benchmark designed to evaluate the fidelity of LVLMs in long-context scenarios by incorporating 8 distinct tasks spanning 6 context length intervals and diverse modalities, including text, images, and videos MMLongCite comprises 8 distinct tasks spanning 6 context length intervals and incorporates diverse modalities, including text, images, and videos. The benchmark is designed to overcome prior limitations by incorporating a wide diversity of data at a larger scale To solve the first shortcoming, it incorporates three visual types-documents, images, and videos-and processes three context formats: image-only, interleaved image-text, and video-only contexts. To solve the second shortcoming, it features 4 broad task categories and 8 distinct longcontext tasks
Context Modalities and Tasks
The benchmark is structured around four distinct tasks: Single-Source Visual Reasoning, Multi-Source Visual Reasoning, Vision Grounding and Video Understanding The construction process is organized into three categories based on the context modality, mirroring the structure presented in Figure 1: image-only, interleaved image-text, and video-only contexts For Image-Only Contexts, five datasets are used: LongDocURL (Deng et al., 2025), MMLongBench-Doc (Ma et al., 2024), HotpotQA (Yang et al., 2018), 2WikiMultihopQA (Ho et al., 2020), and Visual Haystack (Wu et al.) For Interleaved Image-Text Contexts, MM-NIAH (Wang et al., 2024) is represented For Video-Only Contexts, the benchmark is built upon Video-MME (Fu et al., 2025) and LongVideoBench (Wu et al., 2024) Furthermore, MMLongCite-Grounding specifically assesses visual grounding capabilities with scalable difficulty In the Easy setting of MMLongCite-Grounding, models are fed a sequence of composite images, where each composite is stitched from 4 source images In the Hard setting, models need to process a single, large composite image stitched from all source images
Evaluation Metrics and Findings
The evaluation methodology leverages a powerful LVLM, GPT-4.1, as an automated judge for all metrics To measure citation quality, three metrics are reported: Citation Recall (CR), Citation Precision (CP), and their harmonic mean, Citation F1 For Correctness (Cor), the LVLM judge assigns a score on a 3-point scale: 1 for Fully Correct, 0.5 for Partially Correct, and 0 for Incorrect The experiments reveal a significant discrepancy between correctness and faithfulness: many models achieve high correctness scores while exhibiting poor citation performance A key finding is the consistent decoupling of answer correctness from citation fidelity, where models often generate correct answers without accurate source attribution
Analysis of Performance Factors
The evaluation reveals several critical insights into model behavior and context utilization 1) We observe that well-trained, medium-sized models like the MiMo-VL series can match or even surpass significantly larger 72B models on specific tasks 2) Experiments with the Qwen series show a clear trend where citation quality improves with increasing model parameters 3) Our results reveal the evidence of model specialization, such as Gemma3-12B achieving a standout citation F1 score in Video Understanding despite otherwise average performance The impact of Reasoning Mode highlights that Chain-of-Thought (CoT) acts as a double-edged sword, enhancing answer correctness and citation precision but inducing a conservative
citation behavior by overlooking some relevant evidence In MMLongCite-Grounding, the most striking observation is a significant decline across all models when transitioning from Easy
to Hard
setting This decline in grounding accuracy is mirrored by a decline in answer correctness, revealing that dense visual information introduces not only grounding challenges but also distractions that impair the comprehension of models The analysis of the position of needles shows a clear lost-in-the-middle
problem for most Qwen2.5-VL variants, where performance across both correctness and citation declines sharply when the target image resides in the central 40-60% of the context depth
Conclusion and Limitations
Collectively, these findings underscore that merely extending the context window is insufficient and highlight the urgent need for future research to focus on improving the core mechanisms for reliable information utilization and attribution in LVLMs The current scope of MMLongCite primarily centers on visual-centric modalities, namely images, videos, and interleaved text Looking ahead, the plan is to extend this benchmark to encompass a broader spectrum of data types with the ultimate goal of establishing a truly pan-modal citation benchmark in a future version The work was supported by the collaborative project titled Research and Exploration of E-commerce Auditing Agents Based on Multimodal Large Models,
jointly conducted by Soochow University and ByteDance
--- Page 1 ---
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models Keyan Zhou, Zecheng Tang, Lingfeng Ming, Guanghao Zhou, Qiguang Chen, Dan Qiao, Zheming Yang, Libo Qin, Minghui Qiu, Juntao Li and Min Zhang Abstract The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows However an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness are predominantly focused on the textonly domain, while multimodal assessments remain limited to short contexts. To bridge this gap, we introduce MMLongCite, a comprehensive benchmark designed to evaluate the fidelity of LVLMs in long-context scenarios MMLongCite comprises 8 distinct tasks spanning 6 context length intervals and incorporates diverse modalities, including text, images, and videos. Our evaluation of state-of-the-art LVLMs reveals their limited faithfulness in handling long multimodal contexts. Furthermore, we provide an in-depth analysis of how context length and the position of crucial content affect the faithfulness of these models
--- Page 2 ---
MMLongCite MMLongCite-Grounding Image-Only Image [17]: Image [18]: Image [19]: Question: What temperature does the green color of the coffee machine represent for the milk? Model Answer: The green color indicates very cold milk with a temperature of up to 8 °C [19] Image-Text Video-Only Passage [4]: Inspired by the RPG’s and adventures of old on the Game Boy Passage [33]: I have been using make-up ever since I was about 15 years old. I started Question: Which of the following images appears in a certain image of the above document? (A) (B) (C) (D) Model Answer: The image of vase that appears within another image in the provided document is C [21] Question: Which item does the man display to the woman? Model Answer: The man holds a book and shows the woman a picture that is inside of it [3][4] Easy Hard Short Range Long Range Patch Order Patch Order Figure 1 Overview of MMLongCite benchmark Figure 1 The primary task evaluates the processing of multimodal long contexts in Image-Only, Image-Text, and Video-Only formats MMLongCite-Grounding assesses visual grounding capabilities with scalable difficulty In the Easy setting, models are fed a sequence of composite images, where each composite is stitched from 4 source images In the Hard setting, models need to process a single, large composite image stitched from all source images
--- Page 3 ---
Tasks Source Context Length Distribution Total 0∼8k 8∼16k 16∼24k 24∼32k 32∼40k 40∼48k Single-Source Visual Reasoning LongDocURL (Deng et al., 2025) Image [7] Image [7] Image [7] Image [7] Image [7] Image [7] 420 MMLongBench-Doc (Ma et al., 2024) Image [30]Image-Only HotpotQA (Yang et al.
Improvements for AI systems
- Bold Header: Citation Fidelity Grounding Capability Enhancement
The improved AI system can reliably ground its output in the provided context
by incorporating a core citation generation task, as MMLongCite is designed to enforce faithfulness by integrating a core citation generation task within extensive multimodal contexts.
This allows the system to produce verifiable responses with citations, addressing the finding that models often generate seemingly correct answers without faithfully grounding them in the provided context.
- Bold Header: Context Length and Positional Sensitivity Tuning
The system will be tuned to mitigate performance degradation with increasing context length by implementing strategies derived from ablation studies, specifically focusing on the lost-in-the-middle
problem where Qwen2.5-VL variants show a decline when the target image resides in the central 40-60% of the context depth.
- Bold Header: Reasoning Mode Trade-off Management
The system can be optimized to balance correctness and citation precision by dynamically adjusting its reasoning mode; specifically, leveraging Chain-of-Thought (CoT) for deliberative reasoning, boosting the quality of the final answer and the precision of its cited evidence,
while mitigating the associated conservative
citation behavior that causes a general decline in citation recall.
- Bold Header: Visual Reasoning Robustness via Composite Image Synthesis
For visual tasks, the system can be improved to handle dense visual inputs by mastering spatial reasoning through MMLongCite-Grounding's Hard setting,
which requires models to synthesize information across spatially distinct visual sources
from a single, large composite image.
- Bold Header: Retrieval-Augmented Generation (RAG) Noise Mitigation
When employing RAG, the system will be improved by recognizing that for highly capable models, the noise and potential inaccuracies from imperfect multimodal retrieval can outweigh the benefits of context reduction,
suggesting a need to develop retrieval mechanisms that avoid introducing irrelevant data
when processing complex visual contexts.
Abstract
The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness in multimodal settings remain limited to short contexts. To bridge this gap, we introduce MMLongCite, the first benchmark evaluating the faithfulness of LCVLMs via multimodal citation generation. MMLongCite features 2,280 examples across 8 tasks and diverse modalities (image, video, interleaved), with context lengths scaled from 16K to 128K tokens. To test spatial localization capabilities of LCVLMs, we also introduce MMLongCite-HR, evaluating fine-grained visual grounding amidst dense pixel spaces. Through extensive benchmarking of cutting-edge LCVLMs, we provide a systematic analysis of current multimodal citation capabilities. Our results reveal a significant discrepancy between answer correctness and citation faithfulness. We also conduct attention pattern investigations and in-depth error analyses to reveal the underlying phenomena of failures in LCVLMs. MMLongCite establishes a rigorous foundation for diagnosing and advancing the faithfulness of LCVLMs. We hope our findings provide meaningful insights to drive further improvements in the long-context capabilities of LCVLMs.
Sources
- Qwen2.5-VL Technical Report
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
- From Text to Pixel: Advancing Long-Context Understanding in MLLMs
- MiMo-VL Technical Report
- Gemma 3 Technical Report
- Kimi-VL Technical Report
- MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
- Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
- LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
- HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
- LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models