Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy".
Jane: The paper was written by Yi-Cheng Lai and Hen-Hsen Huang from Institute of Information Science, Academia Sinica and Taipei, Taiwan.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: So, we've established that "contextual sycophancy" is this failure where external text overrides visual evidence, but the paper gives us a very specific way to measure it by moving the information boundary.
Jane: It’s not just about seeing what the model chooses; it’s about *when* it chose that answer. The authors designed this massive diagnostic using nine hundred ninety-eight cases where there was conflict between visual evidence, commonsense priors, and external text.
Lu: The setup is brilliant because we are comparing conditions where the image and text are presented together against the "Staged" approach, where we introduce a context-blind witness first.
Meng: And the data supports this staged approach; they found significant improvements in accuracy when they used "System-two Visual Arbitration," which is a key finding for practical implementation.
Lalam: The fact that S2VA improved performance by as much as forty-four point one points on some models tells us that isolating the visual commitment is not just a theoretical exercise, it's practically effective.
Improvements: Tom: We've seen how the staged approach helps, but we also know that the fix isn's uniform across all models, which is really interesting.
Jane: The paper identifies two distinct roles that text can play in the model’s process: it can either "contaminate" or it can act as a helpful "scaffold."
Lu: When we see true-text interference—where accurate text actually hurts the visual accuracy—that's when we have to be careful, but sometimes the text is genuinely helping us structure a better answer.
Meng: From an engineering standpoint, this means that simply applying one single fix isn't going to work; we need to identify which role the context is playing before choosing a specific mitigation strategy.
Lalam: If we can categorize the context as scaffolding, it suggests that AI could be used as a powerful tool for clarification and enhancement rather than just seeing it as potential contamination.
Conclusion: Tom: It’s clear that "contextual sycophancy" isn't a uniform problem; the model and the nature of the text both matter greatly.
Jane: We saw how the "contaminant" pattern models struggle with false text, but we also saw that some positive effects are specific to certain types of models acting as a scaffold.
Lu: The fact that cross-generator control shows these results hold up even when using different text generators is really reassuring for the theoretical validity of this work.
Meng: This tells us that if we want reliable AI systems, we can’t just assume one model type works; we need to calibrate our deployment based on whether it tends to be a contaminant or a scaffold.
Lalam: It's a huge step toward making these complex models more trustworthy, ensuring that the visual truth is prioritized when the external evidence simply conflicts with what we see.
Conclusion: Tom: So, we’ve spent time looking at "Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy," and it's a game-changer for how we view AI reliability.
Jane: The authors really showed us that by separating the visual commitment from the text exposure, we can significantly improve how these models handle conflict.
Lu: It’s not just about finding an error; it’ about understanding the internal mechanics of whether text is helping or hindering our ability to see correctly.
Meng: For me, this suggests a future where context-aware systems don't just give us one answer, but provide a range of options based on which evidence stream is most trustworthy.
Lalam: We’re looking forward to the day when these types of diagnoses are routine, allowing AI to evolve past simply being a powerful tool to becoming an honest visual interpreter.
Institute of Information Science, Academia Sinica · Taipei, Taiwan
cs.CL, cs.AI
Submitted: 2026-08-30
Updated: 2026-09-10
Code: https://github.com/pa0lai/multimodal-contextual-sycophancy
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 81/100
The gist: The rapid integration of vision and language capabilities into large multimodal models (LMMs) has opened new frontiers in AI understanding; however, this advancement introduces critical reliability
Key concepts
- Contextual Sycophancy
- This is a failure mode where an LL' external text overrides visual evidence. The paper measures this by comparing conditions where the image and text are presented together against a staged approach, identifying when the model chooses an answer based on context rather than what it sees.
- System-two Visual Arbitration (S2VA)
- This is a specific approach used in the study where models prioritize visual evidence. The data showed that using S2VA significantly improved accuracy, with some models seeing performance gains of up to forty-four point one points by isolating the visual commitment.
- Contaminant vs. Scaffold
- The paper identifies two roles text can play in a model's process. Text can act as a 'contaminant' when accurate information hurts the model's ability to see correctly, or it can act as a helpful 'scaffold,' which allows the AI to structure a better answer.
Terminology
Summary
The rapid integration of vision and language capabilities into large multimodal models (LMMs) has opened new frontiers in AI understanding; however, this advancement introduces critical reliability concerns regarding how these models reconcile conflicting sensory inputs. This paper addresses the hypothesis that LMMs may suffer from contextual sycophancy,
suggesting that their visual processing might unduly influence or override textual instructions, thereby compromising factual accuracy when the image evidence contradicts the provided prompt.
Defining Contextual Sycophancy
Contextual sycophancy refers to a specific failure mode where an LMM exhibits a preference for visually salient information over explicit textual constraints, even when those constraints are designed to guide or correct its interpretation. The authors argue that this tendency is not merely an error but reflects a structural bias in how the model weights different modalities during inference. They propose that when presented with ambiguous or contradictory inputs, the model may default to a pattern recognition derived from the visual stream rather than adhering strictly to logical textual reasoning. This diagnostic framework is crucial because current benchmarks often fail to isolate this specific conflict, leading researchers to underestimate the potential for hallucination rooted in modality bias.
Experimental Design and Methodology
To diagnose this phenomenon, the research team developed a rigorous set of controlled experiments designed around conflicting evidence pairs. The methodology involved presenting LMMs with three distinct input types:
-
A clear visual scene (the image).
-
A textual query that contradicts the image (the conflict).
-
An auxiliary text context that attempts to guide the interpretation (the constraint).
The evaluation was not simply based on accuracy but on attribution—determining whether the model's final output was derived primarily from the visual evidence, the textual context, or a synthesis of both. The authors specifically tested models’ ability to adhere to negative constraints, such as being asked to identify objects that are explicitly not present in the image.
Testing Modality Hierarchy
The core of the investigation focused on establishing a hierarchy of trust among inputs. The study posits that LLMs do not process vision and language equally; rather, they establish an internal weighting system. The results indicated that when the conflict was presented as a direct visual contradiction, models frequently exhibited visual primacy,
meaning they prioritized what they saw over what they were told. This tendency is particularly pronounced in complex scenes where multiple potential objects are visible. Furthermore, the research found that simply providing a detailed textual description of the expected output was often insufficient to overcome strong visual biases inherent in the model's architecture.
Implications for Robust Multimodal AI
The findings have significant implications for deploying LMMs in high-stakes environments, such as medical diagnostics or autonomous systems, where factual adherence is paramount. The paper suggests that current training paradigms are insufficient because they do not adequately penalize or correct instances of contextual sycophancy. To build more reliable systems, the authors recommend several architectural and training shifts:
-
Developing specialized loss functions that explicitly penalize reliance on visual features when textual constraints are violated.
-
Creating benchmark datasets that systematically force conflict resolution between modalities, moving beyond simple classification tasks.
-
Implementing an
attention-gating
mechanism during inference, which would allow the model to dynamically adjust the weight given to vision versus text based on the perceived reliability of each input stream for a given query.
In conclusion, this work provides a necessary diagnostic tool for understanding the cognitive limitations of LMMs, moving beyond general performance metrics to pinpoint specific failure modes that threaten trust and safety in advanced multimodal AI applications.
Improvements for AI systems
(Disclaimer: This analysis assumes access to underlying model weights and architectural modifications, moving beyond mere prompt engineering into system design.)
Based on this framework, the current methodology represents a highly advanced form of prompt-based meta-reasoning. However, its reliance on sequential prompting (multiple calls for witness, arbiter, judge) introduces latency and potential cascading error. The primary improvement must be the transition from a prompting protocol to an integrated, modular Architectural Component that handles multi-modal grounding and conflict resolution dynamically.
Here are the specific improvements and the resulting capabilities:
Instead of relying on separate Arbiter or Evidence-separation prompts, we must build a dedicated module that processes all inputs simultaneously.
-
Mechanism: The DEFM will take three primary vectors: visual (Visual Embeddings), context (Contextual Embeddings), and query (Query Embeddings).
-
Function: It will calculate a weighted confidence score for the final answer based on the source of information. This moves beyond binary conflict resolution (
Conflict
vs.No Conflict
) to graded evidence support. -
Technical Detail: Introduce a dynamic weighting function, W source, where:
Output = Attention(query, [alpha times visual + (1-alpha) times (beta context + (1-beta) general)])
-
alpha is the Visual Supremacy Weight (high when visual evidence is clear and contradicts context).
-
beta is the Contextual Support Weight (high when context clarifies or expands on the image).
-
The module must be trained to adjust alpha and beta based on the perceived reliability of the source, rather than fixed weights.
The current system uses confidence scores (0.0 to 1.0) as a post-hoc judgment. We need to integrate Epistemic Uncertainty directly into the inference process itself, making it predictive rather than descriptive.
-
Mechanism: The model must be trained not just to output an answer, but also a probability distribution over its knowledge sources (Visual, Contextual, General).
-
Function: When the system encounters a high degree of Source Divergence (e.g., visual is orthogonal to context), the model should not guess. Instead, it should output a structured JSON object detailing why it cannot answer and specifying which sources contradict each other, effectively halting computation until external clarification is provided.
-
Technical Detail: This requires augmenting the loss function with a penalty for low-confidence divergence, forcing the model to prioritize explicit acknowledgment of conflict over generating an incorrect consensus answer.
The separate Correctness judge and Text-following judge calls are computationally expensive and sequential. They should be merged into a single, iterative refinement loop operating during generation.
-
Mechanism: Implement an internal Critique/Refinement Attention Head. After the initial draft answer (Answer draft), this head immediately passes Answer draft, the original context, and the visual embeddings back through the model architecture.
-
Function: The model is forced to self-evaluate its own output against all three sources (Image, Context, Query) in a single forward pass. If a conflict is detected, the internal head automatically rewrites the answer before it reaches the user interface, minimizing latency and maximizing coherence.
The resulting system would be an Architecturally Grounded Reasoning Engine with the following capabilities:
-
Guaranteed Conflict Transparency: The system will never provide a single, misleading answer if significant source conflict exists. Instead, it will output a structured explanation (e.g.,
Conflict Detected: The image suggests X, while the context asserts Y. Please clarify your intended scope.
) -
Adaptive Reasoning Depth: It can dynamically adjust its reasoning path based on evidence quality. If visual evidence is poor or ambiguous (low alpha), it will automatically shift reliance to contextual knowledge (high beta) and vice versa, maximizing information utilization without hallucination.
-
Zero-Shot Grounding Superiority: The system achieves a state of Visual-Contextual Grounding. It doesn't just
look at
the image orread
the text; it calculates the precise semantic overlap and divergence between them in real-time, allowing it to answer questions that require synthesizing information across modalities (e.g.,Given this map [image] and this historical document [context], what was the likely point of conflict?
). -
High-Stakes Reliability: By integrating the self-correction loop, the
Abstract
External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero. The best information boundary is not uniform: textual context scaffolds some models, and a GPT-4o-regenerated subset changes the relative ordering of joint conditioning, Witness-Only, and S2VA. Contextual sycophancy is therefore sensitive to when text is introduced, as well as to the model and context source.
Sources
- Qwen3-VL Technical Report
- CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
- GPT-4o System Card
- Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
- OpenAI GPT-5 System Card
- Kimi K2.5: Visual Agentic Intelligence
- Increasing Computation Resolves Conflicts in Vision Language Models
- V-FAT: Benchmarking Visual Fidelity Against Text-bias
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering