Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP
Nikolette Pedersen, Regitze Sydendal, Veronika Cheplygina, Théo Sourget
IT University of Copenhagen
cs.CV, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 11 pages, 3 figure, poster presentation at the joint FAIMI, BRIDGE, and EPIMI workshop at MICCAI 2026 (Strasbourg, France) conference
Code: https://github.com/nikodice4/MedCLIP_shortcuts
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper investigates how real-world shortcuts manifest across different layers of the medical CLIP-based model MedCLIP and its vision encoder, a frozen ResNet-50.
Terminology
Summary
The paper investigates how real-world shortcuts manifest across different layers of the medical CLIP-based model MedCLIP and its vision encoder, a frozen ResNet-50. The authors attach 17 linear classification probes to the intermediate layers of the ResNet-50 and train them on three different dataset configurations and targets: NIH-CXR14 (pneumothorax) and PadChest (cardiomegaly and pneumothorax). This setup allows them to observe model behaviour during evaluation using subgroup-based calibration and layer-wise confidence curves.
The authors find that the final linear probes achieve a high AUROC but poor calibration in the models. The layer-wise confidence analyses suggest that shortcuts emerge at different depths. Patterns consistent with localised shortcuts, such as drains, appear at later layers, while patterns consistent with diffuse shortcuts, such as scanner-specific noise patterns, emerge earlier, aligning with previous work. Finally, they conduct a manual analysis of the images, which reveals data quality issues in both NIH-CXR14 and PadChest. Their findings underscore that even SOTA models remain vulnerable to shortcuts, and the need for high-quality and well-annotated datasets to draw solid conclusions.
Specifically, the results show that pneumothorax on NIH-CXR14 achieves a global AUROC score of 0.839, with a small gap of 0.028 between drains (0.840) and no drains (0.812). Cardiomegaly on PadChest achieves the highest global AUROC score across all three classification tasks, at 0.905, with the subgroup IDC achieving the highest AUROC score across all dataset configurations and subgroups, reaching 0.917, while its counterpart PMS reaches 0.893. Pneumothorax on PadChest achieves a global AUROC score of 0.875, with PMS and IDC having AUROC scores of 0.801 and 0.852, respectively.
Regarding calibration, the authors find that every model is miscalibrated to some degree, and the number of positive cases for the dataset configurations visibly influences the curves. Pneumothorax on NIH-CXR14, split by whether a patient has a drain, shows the model is poorly calibrated and overconfident, though marginally better calibrated for the drain subgroup. Cardiomegaly on PadChest is the best calibrated model, but still overconfident. Pneumothorax on PadChest is the most miscalibrated model, with the IDC subgroup being the worst calibrated.
For the confidence curves, pneumothorax on NIH-CXR14, split by drain status, shows all four curves stay low and stable across the first 13 building blocks, then diverge from block 13, where positive cases reach high confidence between 0.35 and 0.41, while negative cases stay below 0.25, with the drain curve remaining higher than the no-drain curve. Cardiomegaly on PadChest, split by X-ray machine, shows all four curves remain low, spiking between building blocks 4 and 6 and again at building block 13, but all lines lie in the same range and yield a relatively low confidence, around 0.23-0.28. Pneumothorax on PadChest shows spikes in the earlier layers, with all lines spiking at building block 3 and afterwards remaining stable in their trajectory.
The manual analysis uncovered several errors: for NIH-CXR14, cases of incorrectly entered patient ages and a greyed-out image; for PadChest, patients with conflicting sexes, duplicate images, duplicate metadata, as well as an X-ray of a skull. Further, the automatically annotated drains of the negative cases of pneumothorax in NIH-CXR14 were not completely reliable, with cases where images had been labelled with no drain
but actually contain a drain, found in 3/5 of the images visually examined.
Improvements for AI systems
Improvements to AI Systems:
-
Layer-Aware Shortcut Detection and Mitigation: Implement a monitoring system that attaches lightweight probes to intermediate layers of vision encoders (e.g., ResNet) during training. This system can detect when shortcuts emerge at specific depths (e.g., diffuse scanner noise in early layers, localized artifacts like drains in later layers) and apply targeted regularization (e.g., dropout or adversarial debiasing) only to those layers, reducing reliance on spurious correlations without harming general feature learning.
-
Subgroup-Calibrated Confidence Scoring: Replace global confidence outputs with subgroup-specific calibration curves (e.g., based on known confounders like drain presence, X-ray machine type, or patient demographics). The improved system will dynamically adjust its confidence thresholds per subgroup, reducing overconfidence in minority subgroups (e.g., pneumothorax without drains) and improving clinical trustworthiness.
-
Data Quality Audit Module: Integrate an automated pre-processing pipeline that flags common dataset errors—such as duplicate images, conflicting metadata (e.g., sex, age), mislabeled negative cases (e.g.,
no drain
but drain visible), or non-chest X-ray content (e.g., skull images). This module can be used to clean training data or to weight samples by reliability, preventing models from learning from corrupted labels. -
Shortcut-Aware Evaluation Metrics: Extend standard AUROC reporting with layer-wise confidence curves and subgroup-specific calibration errors (e.g., expected calibration error per subgroup). The improved system will provide a diagnostic report that highlights where shortcuts are most likely (early vs. late layers) and which subgroups are most affected, enabling developers to make informed decisions about model deployment or retraining.
-
Confounder-Aware Training Objective: Modify the loss function to penalize high-confidence predictions that are inconsistent across subgroups (e.g., using a variance penalty on subgroup confidence). This encourages the model to rely on clinically meaningful features rather than machine-specific or annotation artifacts, leading to more robust generalization across hospitals and scanners.
What the Improved AI System Can Do:
-
Detect and localize shortcuts in real time during training, allowing for early intervention before deployment.
-
Provide calibrated, subgroup-aware predictions for medical imaging tasks, reducing false confidence in high-risk subgroups (e.g., pneumothorax with drains) and improving decision support for radiologists.
-
Automatically flag and filter unreliable training data, leading to higher-quality models and more reproducible research findings.
-
Generate interpretable diagnostic reports that show which layers and subgroups are vulnerable, helping clinicians and engineers understand model limitations.
-
Generalize better to unseen hospital data by actively suppressing reliance on scanner-specific noise and annotation artifacts, as demonstrated by the paper’s findings on diffuse vs. localized shortcuts.
Abstract
Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence. However, recent work reveals that CLIP-based models remain vulnerable to shortcuts. We investigate how real-world shortcuts manifest across different layers of the medical CLIP-based model, MedCLIP, and its vision encoder, a frozen ResNet-50. We attach 17 linear classification probes to the intermediate layers of the ResNet-50 and train them on three different dataset configurations and targets: NIH-CXR14 (pneumothorax) and PadChest (cardiomegaly and pneumothorax). This setup allows us to observe model behaviour during evaluation using subgroup-based calibration and layer-wise confidence curves. We find that the final linear probes achieve a high AUROC but poor calibration in the models. The layer-wise confidence analyses suggest that shortcuts emerge at different depths. Patterns consistent with localised shortcuts, such as drains, appear at later layers, while patterns consistent with diffuse shortcuts, such as scanner-specific noise patterns, emerge earlier, aligning with previous work. Finally, we conduct a manual analysis of the images, which reveals data quality issues in both NIH-CXR14 and PadChest. Our findings underscore that even SOTA models remain vulnerable to shortcuts, and the need for high-quality and well-annotated datasets to draw solid conclusions. Code can be found on our GitHub: https://github.com/nikodice4/MedCLIP shortcuts.
Sources
- Limitations of Public Chest Radiography Datasets for Artificial Intelligence: Label Quality, Domain Shift, Bias and Evaluation Challenges
- Dataset Diversity Metrics and Impact on Classification Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models