Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

arXiv:2608.10522 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Yingsheng Liu, Haiming Li, Jingmin Zhu, Jiajun Sun, Victoria Mar, Monika Janda, H. Peter Soyer, Zongyuan Ge, Zhen Yu

Monash University · The University of Queensland

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: INTERNATIONAL CONFERENCE ON MEDICAL IMAGE COMPUTING AND COMPUTER ASSISTED INTERVENTION (ORAL presentation)

Code: https://github.com/Ethan-ysliu/AID

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: The paper "Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training" proposes a novel framework called AID (Adaptive Importance-guided Discretized reconstruction) to

Terminology

Summary

The paper Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training proposes a novel framework called AID (Adaptive Importance-guided Discretized reconstruction) to improve multimodal representation learning for medical image-tabular data. The authors argue that existing multimodal pre-training methods are semantic-agnostic because they treat tabular inputs as flat vectors and use unstable continuous regression objectives. They posit that medical tabular data has a two-dimensional hierarchical structure: at the inter-feature level, clinical attributes have varying diagnostic importance; at the intra-feature level, values within a feature exhibit a continuity-discreteness duality where precise numerical differences often map to shared diagnostic concepts.

To address the inter-feature hierarchy, the framework introduces Importance-Aware Adaptive Masking, which extracts data-driven feature significance offline using principal component analysis and a frozen meta-learning prior (TabPFN v2) to generate a normalized importance vector. This vector modulates the masking rate per feature via the formula rj = min(rbase + αsj, rmax), creating a label-free curriculum that prioritizes diagnostically salient features.

To address the intra-feature duality, the framework introduces a Soft-Label Discretized Module that replaces unstable numerical regression with stable distribution matching. This module uses quantile boundaries to create discrete bins and a triangle kernel to allocate probability mass across adjacent bins, mathematically preserving ordinal relationships of clinical measurements.

The pre-training objective is a composite loss LAID = (LITC + LITM + LDR)/3, combining Image-Tabular Contrastive (ITC), Image-Tabular Matching (ITM) with hard negative mining, and Discretized Reconstruction (DR) losses. The DR loss uses cross-entropy for categorical features and KL divergence for continuous features.

Experiments were conducted on large-scale dermatology datasets (SLICE-3D with over 400,000 images, and HOP with 208,540 images) and an ophthalmology dataset (EyePACS with 88,702 images). The results establish a new state-of-the-art: on the SLICE-3D OOD test set, AID achieves a fine-tuning AUC of 0.944 and pAUC of 0.162, significantly outperforming the strongest baseline TIP (AUC 0.911). On the ID SLICE-3D evaluation, the linear probe AUC reaches 0.984, surpassing TIP's full fine-tuning AUC of 0.971. On HOP, the fine-tuning AUC is 0.926, and on EyePACS, the model achieves an accuracy of 0.740 and QWK of 0.758.

Ablation studies show that removing adaptive masking decreases OOD AUC from 0.944 to 0.930, replacing soft-label with continuous regression yields 0.939, equal-width discretization degrades to 0.914, hard-label discretization recovers to 0.939, and Gaussian kernel achieves 0.940, while the proposed triangle kernel soft-label achieves the peak 0.944. Qualitative analysis with t-SNE projections shows the soft-label objective transforms the latent space into highly ordered, distinct medical semantic clusters, and self-attention heatmaps reveal the model learns a highly structured and importance-aware attention allocation mechanism.

Improvements for AI systems

Improvements to AI Systems:

  1. Hierarchical Feature-Aware Masking for Structured Data: Implement the Importance-Aware Adaptive Masking mechanism in any multimodal or single-modal model that ingests tabular data (e.g., EHRs, sensor logs). Instead of uniform random masking, the AI system will dynamically prioritize masking features with higher diagnostic or predictive significance (derived from PCA and meta-learned priors). This creates a self-supervised curriculum that forces the model to learn robust representations of critical clinical attributes first, improving performance on rare or high-stakes features.

  2. Soft-Label Discretization for Continuous Regression Stability: Replace standard L2/MSE regression objectives for continuous values with the Soft-Label Discretized Module. The AI system will quantize continuous measurements into ordinal bins and assign soft probability distributions (via triangle kernels) across adjacent bins. This eliminates unstable gradient updates from outlier values and preserves the ordinal semantics of clinical measurements. The result is a model that learns smoother, more generalizable decision boundaries, particularly for noisy or skewed continuous inputs.

  3. Semantic-Aware Multimodal Alignment: Integrate the composite loss (ITC + ITM + DR) into vision-language or vision-tabular models. The AI system will align image features with discretized tabular representations rather than raw continuous vectors. This forces the model to map images to shared diagnostic concepts (e.g., mild, moderate, severe) rather than exact numerical values, improving robustness to measurement noise and enabling zero-shot transfer to new clinical sites with different value distributions.

  4. Label-Free Importance Curriculum for Out-of-Distribution Generalization: Adopt the offline importance vector computation (PCA + frozen TabPFN v2) to create a static, data-driven feature ranking. The improved AI system will use this ranking to modulate not only masking but also loss weighting and data augmentation. This leads to models that focus learning capacity on features that remain discriminative across domains, directly improving OOD AUC (as demonstrated: 0.944 vs. 0.911 baseline).

  5. Stable Distribution-Matching Reconstruction for Missing Data: Use the discretized reconstruction loss (KL divergence for continuous, cross-entropy for categorical) as a pre-training objective for tabular encoders. The improved AI system will learn to reconstruct missing or corrupted tabular inputs by predicting probability distributions over discretized bins, rather than exact values. This yields a more robust imputation mechanism that can handle high missingness rates and produces calibrated uncertainty estimates for downstream clinical decision support.

What the Improved AI System Can Do:

  • Diagnose skin lesions and retinal diseases from images with higher accuracy and better calibration, even on unseen patient populations or imaging devices (OOD AUC 0.944, linear probe AUC 0.984 on SLICE-3D).

  • Learn from large-scale medical image-tabular datasets without manual annotation, using self-supervised pre-training that automatically identifies and prioritizes diagnostically critical clinical attributes.

  • Generate stable and interpretable embeddings of continuous clinical measurements (e.g., lesion diameter, retinal thickness) that cluster into distinct medical semantic groups, enabling clearer visualization and clinician trust.

  • Transfer to new medical domains (e.g., from dermatology to ophthalmology) with minimal fine-tuning, achieving strong performance (EyePACS accuracy 0.740, QWK 0.758) without task-specific architectural changes.

  • Provide robust predictions under noisy or incomplete tabular data by reconstructing missing values as probability distributions, reducing the risk of erroneous decisions in real-world clinical workflows.

Sources

Related papers