CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification

arXiv:2604.24201 · cs.LG, q-bio.GN, q-bio.MN · Submitted 2026-04-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification".

Jane: The gist:

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we’ve covered the high-level concept of CMGL, which is essentially using a two-stage approach to tackle multi-omics cancer subtyping by first assessing modality reliability.

Jane: That assessment happens in Stage one through evidential deep learning, where they build confidence scores for each individual data type before any complex reasoning begins <ref:2604.24201#pg1>.

Tom: The core thesis is that existing graph methods fail because they optimize weights jointly with the classification objective, meaning low-quality omics distort the patient similarity graphs and amplify noise during message passing.

Jane: CMGL claims that if you identify which modalities are trustworthy first, then you can perform fusion and graph reasoning under those frozen confidence scores.

Tom: The paper claims this leads to a more robust model because it prevents noisy data from corrupting the features learned by the network in Stage two <ref:2604.24201#pg1>.

Lu: What’s really exciting is that they aren't just doing reliability estimation arbitrarily; they are using a structured evidence modeling approach, constructing Dirichlet parameters to get predictive means and epistemic uncertainty.

Meng: That structural part is what makes it sound more rigorous than just throwing some attention mechanism at the data without a clear way to quantify its quality.

Lalam: If we think about how this helps culture, it means we can build AI systems for cancer subtyping that are inherently more trustworthy because they have built-in quality checks.

Tom: The authors use the MLOmics benchmark, which includes four single-cancer tasks like BRCA and GBM, plus a pan-cancer task to test the generalizability of their framework.

Jane: They also introduce KIRC as an unlabeled cross-cancer transfer cohort for forward inference and prognostic stratification, testing how well the model generalizes its learned knowledge.

Tom: The results show strong performance across these benchmarks, specifically noting an average accuracy improvement of four point zero three percent on the four single-cancer tasks alone <ref:2604.24201#pg1>.

Jane: And on BRCA, they successfully recovered known molecular subtypes and even identified a new metabolic subtype related to tryptophan–kynurenine catabolism.

Tom: But the real impact, as we saw with KIRC, is that a BRCA-trained model transferred without any fine-tuning to stratify KIRC patients into groups matching established ccA/ccB molecular subtypes.

Jane: It suggests that the learned biological axes are not dataset artifacts but are actually conserved structures in cancer biology.

Conclusion: Tom: So, looking at the paper "CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification," we see a framework that separates quality estimation from classification.

Jane: The authors propose this two-stage design to estimate per-sample modality reliability through evidential deep learning, which is then frozen before Stage two <ref:2604.24201#pg1,per-sample modality reliability through evidential deep learning>.

Tom: The key implication is that by filtering out the noise early, you stop low-quality omics from distorting patient similarity graphs and message passing in the graph reasoning stage.

Jane: It means we get a classification result that’t more trustworthy because it’s built on a foundation of assessed data quality.

Tom: The paper shows they hit high accuracy across multiple benchmarks, including single-cancer tasks and the pan-cancer task, and the BRCA model’s transfer to KIRC is quite telling.

Jane: It confirms that this method can capture conserved biological principles across different cancer types because those principles aren't just specific to one dataset.

Tom: Overall, CMGL suggests a practical way to build multi-omics models where data quality is prioritized before feature fusion and graph construction.

Lu: I think the future work should explore how this confidence score can be used for dynamic adaptation if the model encounters novel data types in real-time.

Meng: From an engineering view, we need to figure out how computationally expensive that initial evidence modeling step is when you scale up to hundreds of modalities.

Lalam: It’s fascinating because it opens the door for building more sophisticated AI systems where quality assurance is baked into the design, which should improve overall system performance.

College of Computer Science, Sichuan University · Xinjiang Technical Institute of Physics and Chemistry, Chinese Academy of Sciences

cs.LG, q-bio.GN, q-bio.MN

Submitted: 2026-04-27

Updated: 2026-10-08

Code: https://github.com/chenzRG/CancerMulti-Omics-Benchmark

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: The gist: CMGL proposes a two-stage framework that estimates per-sample modality reliability through evidential deep learning and uses these frozen confidence scores to guide cross-omics fusion and

Key concepts

Evidence Modeling (Stage 1)
This stage uses an encoder to map raw data into a hidden space and an evidence head to predict class-wise evidence. It calculates uncertainty and extracts signals like evidence strength, entropy, and maximum probability to score the trustworthiness of each omics modality.
Modality Confidence Score (r(m)i)
A per-modality scoring network combines three complementary signals—log evidence strength, normalized predictive entropy, and maximum predictive probability—to yield a confidence score for every input modality. This score quantifies how reliable each specific omics data source is for the patient.
Confidence-Guided Fusion
In Stage 2, feature fusion combines omics representations using the frozen confidence scores. Low-confidence modalities are suppressed at both levels of fusion, preventing noisy or unreliable data from diluting the discriminative features learned from trustworthy modalities.
Consistency Intersection Graph (E∩)
Instead of building separate graphs for each modality, CMGL constructs a graph where an edge exists only if it is supported by *all* modalities. This ensures that message passing in the final graph reasoning step is only influenced by consistent, high-quality information.

Terminology

Summary

The gist: CMGL proposes a two-stage framework that estimates per-sample modality reliability through evidential deep learning and uses these frozen confidence scores to guide cross-omics fusion and graph construction for cancer subtype classification.

Motivation

Multi-omics integration can improve cancer subtyping, but modality informativeness and noise vary across cancer types and patients Existing graph-based methods optimize modality weights jointly with the classification objective and therefore lack independent reliability estimates low-quality omics distort patient similarity graphs and amplify noise through message passing Biologically, this means that the molecular signatures of true cancer subtypes, such as the hormone-receptor versus proliferation axis in BRCA, can be diluted by noise from uninformative modalities. The core question is therefore how to identify which omics modalities are trustworthy for each patient before fusion and graph reasoning.

CMGL Framework

CMGL adopts a two-stage design where Stage 1 estimates modality reliability and Stage 2 performs cross-omics feature fusion and graph-based classification under frozen confidence priors The key design principle is that reliability estimation precedes fusion and graph reasoning, and is not updated by downstream gradients.

Stage 1 involves Evidence Modeling where an encoder maps raw input into a hidden space, and an evidence head outputs non-negative class-wise evidence e(m) Following the EDL framework [Sensoy et al., 2018], Dirichlet parameters are constructed to yield the Dirichlet predictive mean π(m)i = E[p(m)i] and the epistemic uncertainty u(m)i = C/S(m). A Quality Estimator extracts three complementary signals: log evidence strength a(m)i, normalized predictive entropy H˜(m)i, and maximum predictive probability π(m)i,max. These are combined by a per-modality scoring network Qm(·) to yield the modality confidence r(m)i. The Stage 1 objective is the weighted sum of three terms: an EDL loss Ledl with an annealed KL regularizer, a confidence-weighted classification loss L(r)cls, and a diversity constraint Ldiv. After training, the confidence vector ri is frozen and used by Stage 2.

Confidence-Guided Fusion and Graph Reasoning

Stage 2 performs feature fusion under frozen confidence priors where independent encoders map each omics modality into a shared representation space, augmented with an omics-identity embedding pm Multi-head self-attention across the M = 4 modality tokens exchanges information feature-wise. The fused representation is calculated as zi = X sum m=1 r(m)i g(m)i⊙ v˜(m) Low-confidence modalities are suppressed at both levels simultaneously, preventing noisy omics from diluting subtypediscriminative features.

For graph construction, CMGL constructs a k-nearest-neighbor (k-NN) graph E(m) independently for each omics modality and retains only edges supported by all modalities: E∩ = M sum m=1 E(m). This ensures that noise from any single low-quality modality cannot reach message passing. The graph message passing uses a two-layer GraphSAGE network [Hamilton et al., 2017] on the consistency intersection graph, with a residual link adding the fused representation zi directly to the output layer.

Results and Interpretability

CMGL consistently improves over baselines on four MLOmics cancer-subtype tasks and the 32-class pan-cancer task, surpassing it by 4.03% in average accuracy on the four single-cancer tasks. On BRCA, the five predicted subtypes correspond closely to the established PAM50 intrinsic subtypes in both molecular signatures and sample proportions. The model also identifies a metabolically active subtype enriched in the tryptophan–kynurenine catabolic pathway. In cross-cancer transfer experiments, the BRCA-trained model transfers without fine-tuning to kidney renal clear cell carcinoma (KIRC), stratifying patients into three prognostically distinct groups whose pathway profiles recapitulate the established ccA/ccB molecular subtype framework. The analysis shows that the proliferation–differentiation axis learned on BRCA transfers as a conserved prognostic axis rather than a dataset artifact.

Conclusion

CMGL estimates modality reliability separately from classification, addressing distinct failure modes by filtering noisy modalities and combining feature-wise signals across modalities. The framework attains the highest accuracy and Macro-F1 on four MLOmics single-cancer tasks and the pan-cancer task. The BRCA embedding recovers the PAM50 subtypes and isolates a metabolic subtype, and the BRCA-trained model transfers to KIRC without fine-tuning, stratifying patients into three clusters matching the ccA/ccB framework. Missing-modality handling and adaptive graph construction are natural extensions of this work.

Data Availability

The datasets used in this study are from the MLOmics benchmark, available at https://github.com/chenzRG/CancerMulti-Omics-Benchmark. Code will be made publicly available upon acceptance.

References

A. R. Brannon, A. Reddy, M. Seiler, A. Arreola, D. T. Moore, R. S. Pruthi, E. M. Wallen, M. E. Nielsen, Huliu Liu and K L Nathanson Molecular stratification of clear cell renal cell carcinoma by consensus clustering reveals distinct subtypes and survival patterns Genes & Cancer 1:152–163 2010.

A. Cheerla and O. Gevaert Deep learning with multimodal representation for pancancer prognosis prediction Bioinformatics 35:i446–i454 2019.

E. Y. Chen, C. M. Tan, Y. Kou, Q. Duan, Z. Wang, G V Meirelles N R Clark and A Ma’ayan Enrichr interactive and collaborative HTML5 gene list enrichment analysis tool BMC Bioinformatics 14:128 2013.

T. Chen, S. Kornblith, M. Norouzi, and G Hinton A simple framework for contrastive learning of visual representations In International Conference on Machine Learning ICML pages 1597–1607 2020.

W. Du, L. Zhang, A Brett-Morris et al HIF drives lipid deposition and cancer in ccRCC via repression of fatty acid metabolism Nature Communications 8:1769 2017.

Y. Gal and Z Ghahramani Dropout as a Bayesian approximation Representing model uncertainty in deep learning In International Conference on Machine Learning ICML pages 1050–1059 2016.

C. Guo, G Pleiss, Y Sun, and K Q Weinberger On calibration of modern neural networks In International Conference on Machine Learning ICML pages 1321–1330 2017.

W. L Hamilton R Ying and J Leskovec Inductive representation learning on large graphs In Advances in Neural Information Processing Systems NeurIPS pages 1024–1034 2017.

Z Han, C Zhang, H Fu, and J T Zhou Trusted multi-view classification with dynamic evidential fusion IEEE Transactions on Pattern Analysis and Machine Intelligence 45:2551–2566 2023.

Y Hasin M Seldin and A Lusis Multi-omics approaches to disease Genome Biology 18:83 2017.

Improvements for AI systems

  1. textbfConfidence-Guided Modality Reliability Estimation in Stage 1: Implementing EDL and Quality Estimator for Robustness Analysis (Section 2.3, Algorithm S1). Developing this capability allows AI systems to quantify modality reliability independently, addressing the issue where low-quality omics distort patient similarity graphs and amplify noise through message passing. This results in a system that can suppress noisy inputs before fusion by using the output of the quality estimator to derive per-modality confidence scores, which are then frozen for Stage 2.

  2. textbfCross-Omics Feature Fusion via Frozen Confidence Gating (Section 2.4.2). The improved system will utilize a mechanism where Low-confidence modalities are suppressed at both levels simultaneously, as described in the text, by using the frozen confidence scores to gate feature aggregation, thereby ensuring that only reliable modalities shape subsequent information flow during fusion and preventing errors from propagating through noisy omics.

  3. textbfConsistency Intersection Graph Construction (Section 2.5). The AI system will employ a method where edges are retained only if every modality independently treats the two patients as neighbors, denoted as E∩ = M m=1 E(m). This ensures that noise from any single low-quality modality cannot reach message passing, leading to a more biologically sound graph structure for patient reasoning.

  4. textbfAdaptive Graph Structure Selection via Warm-up Procedure (Section 2.5.2). Instead of fixed parameters, the system will implement an adaptive neighbor count selection by running a short Stage-2 warm-up in which training updates use the train-only graph and validation Macro-F1 is scored on the train–validation graph. This allows the system to select k∗ = arg maxk F1val(k) dynamically, optimizing graph density for each fold.

  5. textbfRobustness via Hyperparameter Sensitivity Analysis (Section 3.6). The AI framework will be designed to be robust by confirming that chosen operating points lie within a broad robustness region by demonstrating that Macro-F1 varies by less than ±0.025 across tested loss weights, thus providing a stable configuration for training.

Abstract

Motivation: Multi-omics integration can improve cancer subtyping, but modality informativeness and noise vary across cancer types and patients. Most graph methods for multi-omics data learn modality contributions within the downstream classification objective, leaving predictive reliability for each patient implicit. As a result, uninformative modalities can weaken the fused representation, while unreliable omics can introduce noisy patient relationships into graph propagation. To address these two problems, we propose CMGL, which produces a separate reliability estimate before fusion and uses consensus patient neighborhoods for graph classification. Results: CMGL estimates modality confidence for each patient through evidential deep learning, fixes these values during fusion across omics, and performs classification on an independently specified consistency graph. On four MLOmics cancer-subtype tasks and the 32-class pan-cancer task, CMGL consistently improves over the strongest baseline, surpassing it by 4.03% in average accuracy on the four single-cancer tasks. Its representations recover the PAM50 intrinsic subtypes of breast invasive carcinoma (BRCA), and the model trained on BRCA transfers without fine tuning to kidney renal clear cell carcinoma (KIRC), stratifying patients into prognostically distinct groups.

Related papers