ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
cs.CV, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Code: https://github.com/criticaldata/MODALENS
License: http://creativecommons.org/licenses/by/4.0/
The gist: A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image.
Terminology
Abstract
A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with question and report fixed. Under an explicit answer instruction, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, a paired increase of 16.7 points (patient-clustered 95 percent CI 15.6 to 17.7), so report availability reduces image-swap sensitivity under this protocol; the original prompt with a lowercase first-token readout gives 4.70 percent against 17.07 percent, and substitutions also move continuous answer scores where the binary prediction does not change. The labels are derived from reports, which limits conclusions about visual correctness; the direction replicates in two further model lineages. Code, the exact prompts and a run record for every number are at https://github.com/criticaldata/MODALENS.
Sources
- Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
- Understanding intermediate layers using linear classifier probes
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
- TorchXRayVision: A library of chest X-ray datasets and models
- Selective Classification for Deep Neural Networks
- Gemma 3 Technical Report
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
- CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison
- MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Evaluating Object Hallucination in Large Vision-Language Models
- Visual Instruction Tuning
- Vision-language models for chest radiography do not always need the image
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Locating and Editing Factual Associations in GPT
- Steering Llama 2 via Contrastive Activation Addition
- Object Hallucination in Image Captioning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models