Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict
cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: Accepted to Findings of EMNLP 2026
Code: https://github.com/OpenBMB/MiniCPM-o
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal large language models (MLLMs) are increasingly provided with contextual evidence in heterogeneous forms: as a text passage, as a rendered image of the same passage, or as both together.
Terminology
Abstract
Multimodal large language models (MLLMs) are increasingly provided with contextual evidence in heterogeneous forms: as a text passage, as a rendered image of the same passage, or as both together. However, it remains unclear how consistently these surface forms are processed, especially when the evidence conflicts with the model's parametric knowledge. We study modality robustness under knowledge conflict across 13 MLLMs and two datasets, and find them far from robust. (1) Contrary to common belief, models favor a context that contradicts parametric knowledge more readily in image form than in text form; (2) when a contradicting text and image are presented together, the preferred modality is essentially arbitrary, varying with input order, model, and dataset. We further demonstrate that this instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks. To alleviate this brittleness, we examine several simple techniques---prompting, steering, supervised fine-tuning (SFT), and direct preference optimization; the majority prove ineffective, whereas SFT achieves moderate success. We therefore call for greater awareness of this inconsistency and argue that it is fundamental, demanding attention at multiple training stages.
Sources
- Qwen2.5-VL Technical Report
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- How Do Vision-Language Models Process Conflicting Information Across Modalities?
- GPT-4o System Card
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
- Challenges in Understanding Modality Conflict in Vision-Language Models
- MMM-Fact: A Multimodal, Multi-Domain Fact-Checking Dataset with Multi-Level Retrieval Difficulty
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model
- When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering