Scaling an Autoregressive Transformer for Single-Cell Generation
cs.LG, cs.AI, q-bio.GN
Submitted: 2026-08-03
Updated: 2026-09-01
Code: https://github.com/haw-ai-i/SATScG
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors of that cell type.
Terminology
Abstract
We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors of that cell type. For this task we characterize both the biological fidelity of the generated gene expression vectors and the scaling behavior of the pretraining loss. The model is a causal transformer paired with a learned quantized VAE tokenizer, trained with a cross-entropy loss. To evaluate the model, we condition it on held-out gene expression vectors of a cell type and generate vectors of gene expression, comparing the resulting distribution over gene expression vectors to the ground truth distribution of that cell type. We study the scaling properties of the proposed architecture by varying the number of trained parameters and the amount of training data. To our knowledge, we find the first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model. Finally, we discuss how this pretrained model could be finetuned for perturbation response prediction.
Sources
- Sequential Modeling Enables Scalable Learning for Large Vision Models
- Training Compute-Optimal Large Language Models
- Scaling Laws for Neural Language Models
- Scaling Laws for Masked-Reconstruction Transformers on Single-Cell Transcriptomics
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- LLaMA: Open and Efficient Foundation Language Models
- PRiMeFlow: Capturing Complex Expression Heterogeneity in Perturbation Response Modelling
- Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells
- REAL: Response Embedding-based Alignment for LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks