BUZZY: Contrastive Scoring to Mitigate Text-Induced Bias in Multimodal Multiple-Choice QA
cs.AI
Submitted: 2026-03-30
Updated: 2026-09-04
Comments: The current manuscript has substantially diverged from the originally submitted work in scope, methodology, experiments, and authorship, such that the original arXiv record no longer accurately represents the work. We are requesting withdrawal pending guidance on handling the substantially distinct manuscript
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively measurable setting for evaluating vision-language models (VLMs).
Terminology
Abstract
Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively measurable setting for evaluating vision-language models (VLMs). However, because the MCQA format incorporates the candidate choices into the input context, it introduces several unintended biases. Previous work has primarily focused on structural biases, such as preferences for certain choices. Instead, we argue that the choices act as textual priors, causing models to favor linguistically plausible options regardless of the visual content. We hypothesize and empirically verify that a model genuinely relies on visual evidence only when its multimodal distribution significantly diverges from its text-only distribution. Based on this observation we propose BUZZY,a training-free decoding method that corrects multimodal predictions by subtracting the text-only distribution. Experiments with five VLMs on five multimodal MCQA benchmarks demonstrate that BUZZY achieves the highest average accuracy among state-of-the-art methods while reducing inference latency by over 28% compared to prior contrastive decoding approaches. Overall, these results suggest that amplifying the visual signal by penalizing text-only preferences is key to efficient and robust multimodal MCQA reasoning.
Sources
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
- MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding
- Contrastive Decoding Improves Reasoning in Large Language Models
- Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
- SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation
- Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection