Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Revisiting Integration of Image and Metadata for DICOM Series Classification".
Jane: The gist: The proposed end-to-end multimodal framework for DICOM series classification jointly models image content and acquisition metadata using cross-modal attention to produce robust series-level representations,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title and who wrote this stuff. It’s "Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning." The authors are Tuan Truong, Melanie Dohmen, Sara Lorio, and Matthias Lenga from Bayer AG.
Jane: That title really tells you what they're doing—they aren't just looking at images or metadata separately; they’re revisiting how to integrate them using cross-attention and a dictionary learning approach to solve the series classification problem.
Lu: The authors are focusing on making sure that when you combine the visual information and the DICOM tags, you get a robust representation of an entire series, not just one slice in isolation.
Meng: What this means practically is that instead of having separate systems for images and metadata, they’re building one system where they interact constantly to make a better decision about what a whole set of scans looks like.
Lalam: This approach is interesting because it directly addresses the problem mentioned in their abstract: heterogeneous slice content, variable series length, and incomplete metadata.
The paper's summary: Tom: So, the summary boils down to them proposing this end-to-end multimodal framework that uses cross-modal attention to fuse image features and metadata features into a series representation.
Jane: They detail three main parts: first, encoding images and metadata with their own modality-aware modules; second, using a sparse encoder for the metadata that doesn't need imputation; and third, using the 2 point 5D visual encoder to manage variable slice lengths <ref:2602.23833#pg1>.
Lu: The key mechanism is this bi-directional cross-modal attention where both image embeddings and metadata embeddings get to influence each other across all slices.
Meng: And they specifically highlight that because of the sparse metadata encoder, they don't need to fill in in the blanks for missing information, which is a huge practical win.
Lalam: They use this learnable dictionary and value-conditioned modulation within the sparse encoder to contextualize observed metadata features by their semantic identity, which makes it very flexible.
The paper's improvements: Tom: The authors point out several specific improvements they made over previous methods. They are focusing on three main contributions: the end-to-end framework with cross-attention, the sparse metadata encoder with FiLM modulation, and the flexible 2 point 5D visual encoder strategy <ref:2602.23833#pg1>.
Jane: The improvement with the 2 point 5D visual encoder is handling variability in series length and image dimensions by operating on equidistant slices, which helps generalize across different scan sizes <ref:2602.23833#pg1,variability in series length and image>.
Lu: The cross-attention mechanism is key here because it enables cross-modal and cross-slice contextualization, allowing each slice representation to look at all others for better understanding.
Meng: From an engineering standpoint, the sparse metadata encoder's ability to work without imputation noise is a big deal when you are dealing with real-world medical data where headers are often inconsistent.
Lalam: That lack of imputation requirement is what makes it robust to data sparsity; unlike other methods that rely on fixed or learnable imputation which can just add extra noise to the signal.
Conclusion: Tom: So, wrapping this up, the authors show that this framework outperforms image-only and metadata-only approaches across both in-domain and out-of-domain evaluations on liver MRI series classification.
Jane: They prove that modeling the interaction between the visual content and the acquisition data directly leads to better series representations than static fusion methods we’ve seen before.
Lu: This paper shows that explicitly modeling metadata sparsity and cross-modal interactions improves robustness for DICOM series classification, which is a significant step forward in handling real medical imaging data.
Meng: The results show that the performance gain compared to baseline models with simple concatenation is about three percentage points, which confirms that learning how the modalities interact is actually beneficial.
Lalam: And they're not just stopping there; they suggest future work could involve confidence-aware fusion and exploring advanced modulation rules beyond FiLM for even better results.
Tom: That’s it for this one. This paper, "Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning," shows us how to build a system that handles the messiness of medical data by making the image and metadata learn from each other.
Bayer AG
eess.IV, cs.CV
Submitted: 2026-02-27
Updated: 2026-10-07
Comments: Early acceptance at MICCAI 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: The gist: The proposed end-to-end multimodal framework for DICOM series classification jointly models image content and acquisition metadata using cross-modal attention to produce robust series-level
Key concepts
- Sparse Metadata Encoder (SME)
- This component processes DICOM metadata by treating it as a set of observed index–value pairs instead of a dense vector. It uses learnable embeddings for each feature and Feature-wise Linear Modulation (FiLM) to contextualize the scalar values based on their semantic feature identity, allowing the model to handle missing data without guessing or filling in gaps.
- Bi-Directional Cross-Modal Attention (BCA)
- This mechanism allows image representations and metadata embeddings to interact deeply. Vision tokens attend to metadata, and vice versa, using multi-head attention with residual connections. This interaction captures complex relationships between visual features and clinical data across different slices simultaneously.
- 2.5D Visual Encoder
- Since DICOM series have variable lengths, this encoder handles the variability by subsampling an equidistant set of slices to a fixed number (S). It then processes these sampled slices through a backbone and uses cross-slice attention over these tokens to capture global contextual dependencies across the entire series.
- Series-Level Representation Aggregation
- The final step combines the slice-level visual embeddings and metadata embeddings into one single representation for the whole series. This is achieved using a learnable weighting function that assigns importance to each slice embedding, ensuring that irrelevant slices or metadata features are downweighted, leading to a robust series classification.
Terminology
Summary
The gist: The proposed end-to-end multimodal framework for DICOM series classification jointly models image content and acquisition metadata using cross-modal attention to produce robust series-level representations, explicitly handling missing metadata without imputation.
How it works
-
Images and metadata are encoded with modality-aware modules and fused using a bi-directional cross-modal attention mechanism (Page 1).
-
Metadata is processed by a sparse, missingness-aware encoder based on learnable feature dictionaries and value-conditioned modulation, which does not require any form of imputation (Page 4).
-
Variability in series length and image data dimensions is handled via a 2.5D visual encoder and attention operating on equidistantly sampled slices (Page 4).
Model Architecture
The framework treats DICOM series as a variable-length set of slices and associated metadata, where from a series of N slices, S equidistant slices are subsampled to yield an image tensor x ∈ R S×H×W and metadata tensor y ∈ R S×F (Page 2). Each selected slice is center-cropped to 224 × 224, z-score normalized, encoded via an image backbone, and projected to a fixed dimension (Page 5). To capture contextual dependencies between slices, cross-slice attention is employed over slice-level tokens (Page 5). This mechanism allows each slice representation to attend to all other sampled slices, enabling global contextualization and producing a final tensor containing the visual representations denoted as V = [v1,..., vS] ∈ R S×dv (Page 5). In parallel, metadata is encoded with the Sparse Metadata Encoder (SME) into fixed-size embeddings that represent observed DICOM metadata features M = [m1,..., mS] ∈ R S×dm (Page 5).
Sparse Metadata Encoder (SME)
The SME models metadata as a set of observed index–value pairs rather than a dense vector, where each feature index f is associated with a learnable embedding ef ∈ R d (Page 4). Feature-wise linear modulation (FiLM) is used to capture interactions between feature identity and its numeric value by predicting modulation parameters (αs,f, βs,f) = gθ [vs,f, ef] (Page 4). This produces a modulated embedding e˜s,f = ef ⊙ (1 + αs,f) + βs,f that contextualizes the scalar value by its semantic feature identity (Page 4). The modulated embeddings are aggregated across observed features using average pooling to yield a fixed-dimensional representation independent of the number of observed attributes (Page 4). Finally, the aggregated vector is refined via a residual MLP and projected yielding the final embedding ms ∈ R dm (Page 4).
Fusion with Bi-Directional Cross-Modal Attention (BCA)
The vision and metadata embedding tensors V ∈ R S×dv and M ∈ R S×dm are linearly projected into a d-dimensional space, i.e. V˜ = VWv, M˜ = MWm with Wv ∈ R dv×d, Wm ∈ R dm×d (Page 5). Cross-modal interactions are modeled via bi-directional multi-head attention (MHA) with residual connections and layer normalization applied to both the image and metadata pathways (Page 5). The modality-specific outputs are concatenated and projected with a learnable linear layer Wf ∈ R 2d×do to a do-dimensional output space, resulting in F = GELU(LN([V′′,M′′]Wf)) ∈ R S×do (Page 5). A learnable MLP weighting function w: R S×do → [0, 1]S is then used to aggregate slice-level embeddings into a single series-level representation z = X Σ s=1 w(F)sFs ∈ R do (Page 5).
Experiments and Results
The proposed method achieved the best performance across all evaluated methods on the Duke Liver MRI dataset, reaching a weighted F1 score of 96.66% ± 1.03% in five-fold cross-validation (Page 7). This outperformed all baselines, including image-only, metadata-only, and multimodal approaches (Page 7). The sparse, missingness-aware metadata encoder provides a simple mechanism to leverage multi-slice metadata without any kind of imputation (Page 5).
The out-of-domain evaluation showed strong performance across all classes on the in-house test split, with sequence type classification remaining strong for T2, DWI, ADC, and Dixon in-phase (Page 8). However, performance decreased for Dixon opposedphase and portal venous contrast phases (Page 8). The ablation study on the number of input slices S confirmed that cross-modal attention benefits from multiple tokens to align image and metadata representations (Page 9). The limitations suggest that for certain categories, cross-institutional concept shifts may remain a limiting factor for specific classes (Page 9).
The proposed strategy improves performance by approximately three percentage points compared to the best concatenation baseline, highlighting the benefit of learned modality interaction over static fusion (Page 7). The results demonstrate that explicitly modeling metadata sparsity and cross-modal interactions improves robustness for DICOM series classification (Page 1).
The in-domain performance of the proposed method was superior to image-only approaches, confirming that visual information provides strong discriminative cues (Page 7). Models with endtoend image–metadata fusion consistently outperformed unimodal baselines, confirming the complementary nature of the two modalities (Page 7). The results demonstrate that explicitly modeling metadata sparsity and cross-modal interactions improves robustness for DICOM series classification (Page 1). The lower performance of fixed or learnable imputation baselines suggests that imputation noise or complexity degrades classification performance (Page 5). When missingness is high or training data is limited, imputation errors dampen gains from multimodal fusion (Page 5). The ablation related to varying slice counts confirms that cross-modal attention is benefiting from multiple tokens to align image and metadata representations and downweight irrelevant information (Page 9). The proposed method addresses practical challenges of DICOM series classification in a unified way (Page 5).
The final integration to a series-level representation is done by learnable pooling (Page 1). This approach uses fixed-length slice sequences to mitigate the unreliability of single-slice decisions while avoiding the complexity of full 3D modeling (Page 9). The paper concludes that potential improvements include confidence-aware fusion and exploring advanced modulation rules beyond FiLM (Page 9). This finding is consistent with prevalent missing or ambiguous header fields that reduce the utility of metadata and shift the discriminative burden to the image representations (Page 9). The paper considers the application of the method to other medical imaging classification tasks as fruitful to further establish its effectiveness (Page 9).
The proposed method is an end-to-end multimodal framework for DICOM series classification that jointly learns from image content and metadata while explicitly accounting for incomplete metadata and all the aforementioned challenges (Page 1).
Improvements for AI systems
- Bold header: Sparse Metadata Encoder (SME) implementation
This module models metadata as a set of observed index–value pairs
and uses a learnable dictionary
combined with value-conditioned modulation.
This allows the system to encode observed metadata without requiring any form of imputation, making it resilient to missing DICOM header data.
- Bold header: Bi-directional Cross-Modal Attention (BCA) fusion
The framework employs a bi-directional multi-head attention (MHA)
mechanism where visual features and metadata are reciprocally modulate each other across slices.
This allows for a richer, contextually aware series representation by enabling cross-modal and cross-slice contextualization.
- Bold header: 2.5D Visual Encoder strategy
The system utilizes a 2.5D visual encoder
where each slice representation to attend to all other sampled slices, enabling the emphasis of relevant content while down-weighting redundant or irrelevant information.
This handles variability in series length and image data dimensions by operating on equidistantly sampled slices.
- Bold header: End-to-end multimodal framework
The proposed system is an end-to-end multimodal framework for DICOM series classification that jointly learns from image content and metadata while explicitly accounting for incomplete metadata and all the aforementioned challenges.
This results in a unified representation that consistently outperforms relevant image-only, metadata-only and multimodal 2D/3D baselines.
- Bold header: Robustness to data sparsity
The SME's design ensures resilience against missing data because it does not require any form of imputation,
unlike other methods that use fixed or learnable imputation
which can introduce noise. This prevents imputation errors dampen gains from multimodal fusion
when missingness is substantial.
- Bold header: Out-of-domain generalization
The model is tested for generalization by training on a large in-house cohort and testing on the Duke dataset, showing that it maintains strong performance across various sequence types, with sequence-type classification remains strong for T2, DWI, ADC, and Dixon in-phase.
Abstract
Automated identification of DICOM image series is essential for large-scale medical image analysis, quality control, protocol harmonization, and reliable downstream processing. However, DICOM series classification remains challenging due to heterogeneous slice content, variable series length, and entirely missing, incomplete or inconsistent DICOM metadata. We propose an end-to-end multimodal framework for DICOM series classification that jointly models image content and acquisition metadata while explicitly accounting for all these challenges. (i) Images and metadata are encoded with modality-aware modules and fused using a bi-directional cross-modal attention mechanism. (ii) Metadata is processed by a sparse, missingness-aware encoder based on learnable feature dictionaries and value-conditioned modulation. By design, the approach does not require any form of imputation. (iii) Variability in series length and image data dimensions is handled via a 2.5D visual encoder and attention operating on equidistantly sampled slices. We evaluate the proposed approach on the publicly available Duke Liver MRI dataset and a large multi-institutional in-house cohort, assessing both in-domain performance and out-of-domain generalization. Across all evaluation settings, the proposed method consistently outperforms relevant image only, metadata-only and multimodal 2D/3D baselines. The results demonstrate that explicitly modeling metadata sparsity and cross-modal interactions improves robustness for DICOM series classification.
Sources
Related papers
- VesselSDF: Distance Field Priors for Vascular Network Reconstruction
- cSVR: Convolutional Slice-to-Volume Reconstruction
- NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
- AneumoBench: A Source-Linked Benchmark for Synthetic-Geometry Transfer in Aneurysm CFD
- RETO: A Rotary-Enhanced Transformer Operator for High-Fidelity Prediction of Automotive Aerodynamics
- SAMRI-2: A Memory-based Model for Cartilage and Meniscus Segmentation in 3D MRIs of the Knee Joint