Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data

arXiv:2608.13256 · cs.LG, cs.AI · Submitted 2026-08-13 · Read on arXiv

Francesca Pia Panaccione, Sofia Mongardi, Marco Masseroli, Pietro Pinoli

Politecnico di Milano

cs.LG, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper "Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data" presents a comparative analysis of generative models for transcriptomic data, investigating strategies to

Terminology

Summary

The paper Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data presents a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs to ensure synthetic data capture real-world gene patterns. The authors introduce and benchmark three variants of Generative Adversarial Networks (GANs), with MK-TGAN (Multi-Kernel Transcriptomic Generative Adversarial Network) standing out for its performance in terms of both realism and utility of generated data.

The paper addresses challenges in biomedical research such as dataset imbalances, biases, and ethical/legal constraints limiting access to high-quality data. The authors note that "the ability to generate synthetic data that accurately replicate the statistical properties of real genomic data could help address several challenges, such as rebalancing datasets (through data augmentation) and enabling their use without the legal and ethical limitations associated with data derived from patients."

The three proposed models are:

  1. Graph-Modulated GAN (GM-GAN): A two-layer dense feed-forward network where prior biological knowledge is incorporated by modulating the learnable parameters of the last layer of the generator using a Hadamard product with the correlation matrix A, so the generator's weights are scaled according to the strength of known biological relationships.

  2. Graph-Regularized GAN (GR-GAN): Shares the same topology as GM-GAN but incorporates prior knowledge indirectly by regularizing the learned weight structure, allowing the generator to remain flexible while still being guided by prior knowledge. The cost function is augmented with two terms: λ1∥W − A∥2 and λ2ΣReLU(−(A ⊙ W)i,ⱼ), where the first enforces structural similarity and the second discourages learned relationships that contradict prior biological knowledge.

  3. MK-TGAN: The most original contribution, featuring multiple GNN kernels connected to non-learnable graphs, where genes are represented as nodes and edge weights reflect prior biological knowledge. The generator employs multiple parallel GNN kernels, each consisting of several graph convolutional layers that process independent embedding representations, with outputs combined through a learnable weight matrix WF. The architectural novelty compared to BioGAN is that MK-TGAN adopts a multi-kernel formulation where multiple independently parameterized GNN kernels operate in parallel on the same biological prior, capturing complementary and heterogeneous biological relationships.

The experimental setup used breast cancer transcriptomic data from TCGA (1,127 samples with subtypes Basal, Her2, LumA, LumB), focusing on 197 EMT-related genes from MSigDB. The biological prior was constructed from 396 healthy breast tissue samples from GTEx, computing pairwise Pearson correlation coefficients to create a co-expression-based knowledge graph. The correlation matrix A spans [−0.741, 0.976] with negative values accounting for 38.6% of off-diagonal elements, and no post-hoc preprocessing was applied.

All models were trained using conditional WGAN-GP framework with gradient penalty (λ=100), 600 epochs, latent space dimension 128, RMSprop optimizer with learning rate 5×10−4, and the critic updated twice per generator update. MK-TGAN used K=4 parallel GNN kernels.

Key results from Table 1 show:

  • Unsupervised metrics: MK-TGAN achieved the best recall (0.784) and highest correlation score (0.914) while maintaining very high precision (0.960). BioGAN showed high precision (0.837) but low recall (0.381), indicating it accurately reproduce[s] a limited subset of the real data, while failing to capture its broader variability.

  • Detectability: MK-TGAN yielded markedly lower detectability (MLP accuracy = 0.760), indicating generated samples are more similar to real data and therefore harder to distinguish. Other methods (GAN, WGAN-GP, BioGAN) exhibited near-perfect detectability (accuracy ≈ 1.0).

  • Utility (TSTR): MK-TGAN consistently achieved the best downstream performance (LR accuracy = 0.805, MLP accuracy = 0.791), and all models incorporating prior knowledge—including BioGAN—outperform the baselines without prior knowledge.

  • Parameter efficiency: MK-TGAN uses fewer than 19,000 learnable parameters, compared to over 100,000 in the other models.

The authors conclude that "(i) incorporating prior biological knowledge improves both the realism and utility of synthetic transcriptomic data compared to unconstrained baselines and implicit regularization strategies that do not rely on graph-structured generation; (ii) single-stream graph-informed generators can achieve high precision but may suffer from limited recall and reduced variability; and (iii) the multi-kernel MK-TGAN formulation provides a more balanced trade-off, increasing diversity (recall) while preserving high fidelity (precision)."

A limitation acknowledged is that experimental validation is restricted to breast cancer transcriptomic data and to a biologically focused set of EMT-related genes, though the architecture can, in principle, be extended to broader transcriptomic domains. Future work will focus on evaluating MK-TGAN on larger gene panels and whole-transcriptome datasets, spanning multiple cancer types and biological processes, and incorporating richer biological priors—such as regulatory, pathway, and protein–protein interaction networks.

Improvements for AI systems

Improvements to AI systems:

  1. Multi-Kernel Graph-Structured Generators: Implement parallel, independently parameterized graph neural network (GNN) kernels that process the same biological prior graph simultaneously, then fuse outputs via a learnable weight matrix. This captures complementary and heterogeneous relationships, increasing output diversity (recall) without sacrificing fidelity (precision), unlike single-stream graph generators that overfit to a narrow data subset.

  2. Graph-Modulated Parameter Scaling: Replace dense-layer weight initialization with a Hadamard product between learnable weights and a precomputed biological correlation matrix (e.g., gene co-expression). This directly scales generator weights according to known biological relationship strengths, embedding domain knowledge at the architectural level rather than as a post-hoc constraint.

  3. Dual-Objective Graph Regularization: Augment the generator loss with two terms: (a) Frobenius norm penalty between learned weights and the biological prior matrix to enforce structural similarity, and (b) a ReLU-based penalty on negative correlations of the Hadamard product to actively discourage learned relationships that contradict prior knowledge. This balances flexibility with biological plausibility.

  4. Conditional WGAN-GP with Graph Priors: Train all generative models under a conditional Wasserstein GAN with gradient penalty (λ=100), using a critic updated twice per generator update, RMSprop (lr=5×10−4), and 600 epochs. This stabilizes training while allowing graph-regularized generators to produce samples with lower detectability (MLP accuracy 0.76 vs 1.0 for unconstrained baselines).

  5. Parameter-Efficient Graph Generation: Design the generator to use fewer than 19,000 learnable parameters (vs >100,000 in baselines) by sharing graph convolution layers across kernels and using non-learnable graph adjacency matrices. This reduces overfitting risk and computational cost while maintaining high utility.

What the improved AI system can do:

  • Generate synthetic transcriptomic data (e.g., gene expression profiles) that accurately replicate real-world gene co-expression patterns, with high precision (>0.95) and recall (>0.78) simultaneously.

  • Rebalance imbalanced biomedical datasets (e.g., rare cancer subtypes) via data augmentation, producing synthetic samples that are statistically indistinguishable from real patient data (detectability 0.76).

  • Enable downstream machine learning tasks (e.g., cancer subtype classification) with performance comparable to training on real data (accuracy 0.80), even when only synthetic data is used for training.

  • Incorporate diverse biological priors (co-expression, regulatory networks, protein-protein interactions) flexibly, either as hard architectural constraints or soft regularizations, depending on data availability and desired trade-off between fidelity and diversity.

  • Scale to larger gene panels and whole-transcriptome datasets across multiple cancer types, with minimal parameter overhead, while maintaining biological realism and utility for clinical and ethical data-sharing applications.

Sources

Related papers