Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions
cs.AI, cs.LG
Submitted: 2026-03-30
Updated: 2026-08-28
Comments: Accepted to EMNLP 2026 Main Conference
Code: https://github.com/long21wt/scaffold-effect
License: http://creativecommons.org/licenses/by/4.0/
The gist: Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts.
Terminology
Abstract
Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on two clinical neuroimaging cohorts for binary classification of affective disorders and cognitive decline. Both cohorts include structural magnetic resonance imaging (MRI) acquired under their original research protocols. Prior work does not establish the included neuroimaging inputs as reliable stand-alone diagnostic evidence for the present tasks. Nevertheless, when neuroimaging context is introduced, smaller VLMs gain up to 0.66 F1 under the evaluated augmented conditions, becoming competitive with models an order of magnitude larger. Confidence estimation shows that most of the calibration improvement for the analyzed smaller models occurs after the MRI reference is added to the prompt, before any image is supplied. Our preliminary expert case study finds that faithfulness remains low in every condition examined, with the reviewed model introducing unverified clinical details. Finally, in our single-model intervention, preference alignment suppresses MRI-referencing behavior but reduces the augmented-condition advantage, leaving the underlying issue unresolved. These results caution against reading surface metric gains as evidence of true multimodal integration, with direct implications for clinical VLM deployment.
Sources
- Qwen3-VL Technical Report
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Ministral 3
- MedGemma Technical Report
- Roleplaying with Structure: Synthetic Therapist-Client Conversation Generation from Questionnaires
- Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen2.5-VL Technical Report
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection