Effective Biological Representation Learning by Masking Gene Expression

arXiv:2605.31562 · cs.LG · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Effective Biological Representation Learning by Masking Gene Expression".

Jane: RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery.

Tom: First, who's behind it and why it matters.

Paper summary: Jane: They are proposing TxFM as their answer, suggesting a masked autoencoder specifically tailored for RNA-seq count data. The thesis here is that by focusing on inductive representation learning, they aim to create representations that can successfully transfer to cellular contexts that the model has never encountered during its initial training phase.

Lu: It’s important to note the context they are operating in; they are addressing the issue of foundation models underperforming relative to linear baselines on held-out tasks like those reported by Szałata et al., two thousand twenty-four and others, which suggests current scaling isn't enough <ref:2605.31562#pg1>.

Meng: So, what makes TxFM different from just another large model? The paper hints that the approach is tailored specifically to the nature of count data, which implies a smarter way of handling the inherent noise in gene expression measurements compared to standard image or language models.

Lalam: For culture, this means we’re looking at a way to build AI systems that don't just memorize patterns but actually learn the underlying biological rules so they can apply those rules to entirely new cellular conditions. That’s a significant cultural shift in how we approach modeling biology.

Tom: Right, and the authors are making this matter because they are showing that their method, TxFM, combined with a specific curated public corpus called DiverseRNA-1 point 4M, can generate high-fidelity gene representations that surpass competing foundation models on held-out genetic perturbation datasets <ref:2605.31562#pg2>.

Jane: So it boils down to this: they developed TxFM via an MAE framework for transcriptomics and argue that its inductive learning capability, when trained on diverse data, provides a superior way to represent gene expression than what current large foundation models achieve.

Lu: The diversity of the training set is key here; they constructed DiverseRNA-1 point 4M which includes a mix of single-cell, bulk data from various tumor types and cell lines like K562, which gives them a broad view of biological variability <ref:2605.31562#pg2>.

Meng: It sounds like the success hinges heavily on that data curation aspect, because if the training data isn't representative enough or if they don't handle gene set variations well, those strong representations might not generalize as much as they seem in the initial reports.

Lalam: If we can get AI to learn robust representations from such a rich, diverse corpus, it means we could start modeling complex disease states with much higher fidelity than we can now.

Conclusion: Jane: I think the main implication is that for transcriptomics, we need to move beyond just building bigger foundation models and start focusing intensely on designing architectures and training objectives that are intrinsically suited to the statistical properties of gene expression counts.

Lu: They’re showing that the specific choices they made—like using a Poisson-based loss function instead of something like Negative Binomial—are not just technical details; they directly impact how well the model learns relationships, suggesting a deeper understanding of the underlying math is required for these models.

Meng: From a practical standpoint, this means when we design new AI tools for drug discovery or personalized medicine based on gene expression, we should prioritize learning methods that are proven to have strong inductive transfer capabilities across different biological samples.

Lalam: This paper gives us the confidence that by synthesizing careful data curation with smart self-supervised modeling, we can build AI systems that capture fundamental biological relationships in a way that is more reliable and applicable to real-world problems.

Tom: It’s a strong call for the community to look closely at how they structure these models and what objectives they optimize for when dealing with high-dimensional count data. The authors are pointing toward future research into mining these pretrained models to discover new mechanisms, biomarkers, and drug targets.

Jane: So, the future direction seems very clear: we need more work on combining this kind of inductive SSL with the curation strategies they used to see how powerful these representations can truly be for finding novel biological insights.

Recursion Research Institute of Technology (Recursion) · Valence Labs

cs.LG

Submitted: 2026-05-29

Updated: 2026-09-28

Code: https://github.com/recursionpharma/opentxfm

Importance score: 89/100

The gist: RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery.

Key concepts

Masked Autoencoder (MAE)
This is a self-supervised learning technique where the model learns to reconstruct missing parts of the input data. In this context, it means the model is shown a partial gene expression vector and must predict the values of the masked genes to recreate the original full vector.
Poisson-based Reconstruction Loss
This is a specific mathematical loss function used during training that models count data, like gene expression. It focuses on reconstructing genes with low-to-moderate expression more accurately than other methods, as it generates gradients proportional to the prediction error relative to the actual counts.
Inductive Representation Learning
This refers to learning gene representations that are general enough to apply successfully to new, unseen biological contexts. The goal is for the model's understanding of a gene's function or relationship, learned from training data, to be useful when applied to entirely different cell types or conditions.
DiverseRNA-1.4M
This is a curated public dataset consisting of 1.4 million bulk and single-cell RNA-seq samples with broad genetic coverage. This extensive data was crucial for training TxFM, allowing it to learn robust representations across many different biological conditions.

Terminology

Summary

RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery. The gist: inductive self-supervised learning (SSL) via a masked autoencoder (TxFM) trained on a curated public corpus outperforms foundation models trained on much larger atlas-scale datasets for generating high-fidelity gene representations for downstream tasks.

Model Architecture and Training Objective

TxFM is a self-supervised learning (SSL) model designed for the masked reconstruction of partially observed gene expression count data, utilizing an asymmetric Masked Autoencoder (MAE) framework. The input, a gene expression count vector, is processed by a transformer encoder that maps unmasked genes to a bottlenecked CLS-token representation. This token is then passed through an MLP decoder to reconstruct the full expression vector. A novel rectified tanh activation function, defined as Equation 1, is introduced in the decoder to constrain outputs to data support while preserving gradient flow, which asymptotically approaches the library size. The model is trained using a Poisson-based reconstruction loss (Equation 2), which is derived from the Poisson negative log-likelihood of count data. This loss function prioritizes reconstruction of genes with low-to-moderate expression, as it generates gradients proportional to the relative prediction error rather than models like MSE or SmoothL1.

Data Curation and Preprocessing

The success of TxFM relies heavily on careful training data curation, specifically the construction of DiverseRNA-1.4M, a public corpus comprising 1.4 million bulk and single-cell RNA-seq samples with broad genetic coverage. A key strategy involves handling varying gene sets across datasets by learning representations for every gene in the vocabulary but only applying reconstruction loss on genes that are actually measured in an individual sample, avoiding zero-padding. Preprocessing involves two standard steps: first, library size normalization (setting the total count to a fixed size, e.g., 105), followed by log-transformation after adding a pseudocount: xi = log(xi + 1), or log1p.

Inductive Representation Learning and Ablation

The work focuses on inductive representation learning, aiming for representations that transfer to cellular contexts unseen during training. The authors systematically investigate crucial architectural configurations through an ablation study. Key findings include:

  1. The Poisson-based loss function significantly improves performance compared to Negative Binomial (NB) or Zero-Inflated NB (ZINB) losses, as the latter can suffer from vanishing gradients for overestimation errors on high-dispersion genes.

  2. The rectified tanh activation function is necessary for strong transfer performance, as pure ReLU activations diverge during training.

  3. The uniform masking strategy (sampling a subset of size K=2048) yields strong performance, despite larger masks offering higher compute costs; the authors also tested frequency-weighted masking based on gene sparsity.

  4. Dataset curation is vital: including phenoprint-curated perturbed K562 cells and bulk RNA-seq data significantly improves perturbation representations, demonstrating a powerful synergy unlocked by using SSL to jointly model single-cell and bulk transcriptomic modalities.

Performance Benchmarks and Biological Discovery

TxFM demonstrates superior performance in inductive SSL settings, achieving the highest overall scores on held-out genetic perturbation datasets compared to 16 competing foundation models. Furthermore, the learned gene representations themselves are valuable:

  1. The decoder gene weights achieve state-of-the-art recovery of known biological relationships without direct supervision.

  2. Gene representation learning in model parameters shows that TxFM obtains the highest recall for whole-genome known gene-gene relationships across various databases, suggesting strong potential for novel relationship discovery in service of identifying new drug targets.

  3. The study confirms that many fundamental biological relationships emerge naturally in lower-rank manifolds of high-dimensional transformer gene embedding spaces.

Inference and Transferability

When evaluated on unseen perturbed cells (inference-time cell perturbation representation learning), TxFM-B trained on DiverseRNA-1.4M achieves the highest overall score, outperforming competing FMs despite having nearly 4× fewer parameters and training on 100× less data. The results indicate that architectural choices are impactful at a fixed data scale, and the advantage is not solely attributable to perturbational overlap with the evaluation data, but rather to the broader curation of DiverseRNA and the model architecture. The study also shows that fine-tuning TxFM-B pre-trained on DiverseRNA1.4M outperforms other models fitted to the same target data, yielding a 5–19% relative improvement.

Conclusion

TxFM is presented as a viable modeling approach for transcriptomics representation, provided there is a careful synthesis of model architecture and training data curation. The work motivates future research into mining pretrained models to discover new mechanisms, biomarkers, and drug targets.

Improvements for AI systems

As a fastidious researcher, I have analyzed the TxFM model and its methodology described in this paper. The core innovation lies in developing a self-supervised masked autoencoder (TxFM) specifically tailored for high-dimensional, noisy RNA-seq count data, emphasizing inductive representation learning over simple reconstruction.

Here are the specific improvements to AI systems and what those improved systems can achieve:


The primary improvements stem from the architecture and training strategy of TxFM, which focuses on learning biologically meaningful gene representations that generalize well to unseen cellular contexts (inductive representation learning).

  1. Improvement: Implement a self-supervised Masked Autoencoder (TxFM) architecture using a Transformer encoder with a learnable CLS token and an MLP decoder.

  2. Improvement: Utilize a domain-specific reconstruction loss based on the Poisson likelihood, stabilized by a novel rectified tanh activation function, instead of standard MSE or SmoothL1 losses.

  3. Improvement: Employ library size normalization followed by log1p transformation for count preprocessing to stabilize comparisons across samples with varying sequencing depths.

  4. Improvement: Train the model on a curated, diverse public corpus (DiverseRNA-1.4M) rather than relying solely on massive, general atlas-scale datasets for initial pretraining.

  5. Improvement: Incorporate sophisticated data curation strategies (e.g., phenoprint curation of perturbed cells) to enrich training data and mitigate potential data leakage issues when using cell lines with similar assay protocols in the evaluation set.

The resulting improved AI system (TxFM) can perform the following specific tasks:

  1. Perform high-fidelity reconstruction of gene expression counts from partially observed (masked) data, effectively learning the underlying biological distribution of gene expression patterns.

  2. Generate low-dimensional, biologically meaningful embeddings for genes that capture functional relationships (e.g., protein complexes or regulatory networks) by analyzing the learned weights in the decoder MLP and codebook representations of the transformer encoder.

  3. Predict and accurately characterize gene expression changes induced by genetic perturbations (e.g., CRISPRi), achieving high scores in metrics like Perturbation Consistency and Linear Separability on unseen cell lines, even without direct supervision on those new perturbations (zero-shot generalization).

  4. Discover novel biological relationships between genes by analyzing the learned gene parameter matrices, which are shown to encode known functional connections (e.g., protein-protein interactions) in a lower-dimensional manifold than many competing models.

  5. Provide robust, batch-corrected cell type clustering and classification, ensuring that learned representations are invariant to technical noise (batch effects) while still preserving the ability to distinguish true biological cell types (high ASW and high classification probing accuracy).

Sources

Related papers