When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models

arXiv:2608.19208 · cs.CL, cs.CV · Submitted 2026-06-12 · Read on arXiv

cs.CL, cs.CV

Submitted: 2026-06-12

Updated: 2026-09-01

Code: https://github.com/Wangyf1998/Irrelevant_Text_Matters

License: http://creativecommons.org/licenses/by/4.0/

The gist: Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored.

Terminology

Abstract

Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant context by formulating it as a controlled intervention within a binary visual judgment framework. By maintaining an invariant prompt structure while varying auxiliary inputs, we observe that irrelevant text consistently biases model predictions across diverse benchmarks. To move beyond performance metrics, we characterize this sensitivity through a decision margin defined by the log-probability difference between binary candidates. Our analysis reveals a robust geometric regularity: contextconditioned margins follow a consistent affine transformation of their context-free counterparts. This finding demonstrates that irrelevant context does not manifest as unstructured stochastic noise but as a estimable distortion of model preference. We further interpret the fitted affine parameters as metrics for visual commitment preservation and directional answer bias. These findings provide a margin-level diagnostic view of irrelevant-context effects in MLLMs and offer a basis for future studies on noisy-context robustness

Sources

Related papers