When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models
cs.CL, cs.CV
Submitted: 2026-06-12
Updated: 2026-09-01
Code: https://github.com/Wangyf1998/Irrelevant_Text_Matters
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored.
Terminology
Abstract
Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant context by formulating it as a controlled intervention within a binary visual judgment framework. By maintaining an invariant prompt structure while varying auxiliary inputs, we observe that irrelevant text consistently biases model predictions across diverse benchmarks. To move beyond performance metrics, we characterize this sensitivity through a decision margin defined by the log-probability difference between binary candidates. Our analysis reveals a robust geometric regularity: contextconditioned margins follow a consistent affine transformation of their context-free counterparts. This finding demonstrates that irrelevant context does not manifest as unstructured stochastic noise but as a estimable distortion of model preference. We further interpret the fitted affine parameters as metrics for visual commitment preservation and directional answer bias. These findings provide a margin-level diagnostic view of irrelevant-context effects in MLLMs and offer a basis for future studies on noisy-context robustness
Sources
- Rich Knowledge Sources Bring Complex Knowledge Conflicts: Recalibrating Models to Reflect Conflicting Evidence
- Microsoft COCO Captions: Data Collection and Evaluation Server
- MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Re-ranking the Context for Multimodal Retrieval Augmented Generation
- Improving Few-Shot Performance of Language Models via Nearest Neighbor Calibration
- Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
- Why Language Models Hallucinate
- Studying Large Language Model Behaviors Under Context-Memory Conflicts With Real Documents
- Reducing Language Biases in Visual Question Answering with Visually-Grounded Question Encoder
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- An Explanation of In-context Learning as Implicit Bayesian Inference
- Corrective Retrieval Augmented Generation
- mR squared AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA
- Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering