TorchGWAS 1.0: GPU-accelerated GWAS at scale

arXiv:2604.21095 · cs.DC, cs.SE, q-bio.GN · Submitted 2026-04-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: Today's paper: "TorchGWAS 1.0: GPU-accelerated GWAS at scale".

Marcus: TorchGWAS is a framework designed for high-throughput association testing across large panels of quantitative traits by leveraging GPU acceleration to overcome computational bottlenecks in phenotype-rich screening workflows.

Ines: First, who's behind it and why it matters.

Paper summary: Ines: We’ve discussed the technical mechanics of TorchGWAS one point zero and what it claims to achieve, and now we need to reflect on the bigger picture regarding its title and who developed this work.

Marcus: Before we wrap up, I think it helps to remember that the authors are focused on creating a framework specifically designed for phenotype-rich studies where the same genotype matrix is reused across a large panel of quantitative traits.

Yuki: Knowing their focus helps me see how they prioritized efficiency in handling those massive genotype matrices over other potential modeling complexities, which speaks volumes about their design philosophy.

Ines: The title itself, "TorchGWAS one point zero: GPU-accelerated GWAS at scale," tells us immediately that the main goal is making genome-wide association studies faster by using GPU acceleration for large scales of quantitative traits.

Marcus: That acceleration is crucial because it directly addresses the computational bottleneck we talked about earlier, allowing for much larger panels to be screened in a practical timeframe.

Yuki: From a population genetics viewpoint, this work suggests that complex trait mapping can become significantly more accessible when we have the computational power to screen many traits concurrently.

Ines: The implication is that researchers in fields like imaging genetics can conduct rapid, large-scale screening of candidate traits much more efficiently than they could before with traditional methods.

Marcus: Exactly, it’s about making high-throughput quantitative trait association a more accessible reality for those dealing with huge amounts of phenotype data from a single cohort.

Yuki: I see this as accelerating the pace at which we can test hypotheses about how genetic architecture influences many different traits in complex organisms.

Ines: It really puts the power back into the hands of the researcher who needs to focus on interpreting those results, rather than spending all their time managing extremely slow computational pipelines.

Conclusion: Ines: So, we've spent some time looking at how TorchGWAS one point zero handles these large phenotype panels, and now we need to wrap up by discussing the title and the people behind this work.

Marcus: Yeah, before we move on to the broader implications of this software, it’s important for us to anchor ourselves in what TorchGWAS one point zero actually is and who developed it.

Yuki: I think knowing the authors helps me get a better sense of their perspective on why they chose this specific approach for managing those huge genotype matrices.

Ines: Exactly, Yuki, because the title itself—"GPU-accelerated GWAS at scale"—is pretty descriptive; what does that actually mean for someone trying to handle phenotype-rich data?

Marcus: From my side, it means they've tackled the core computational bottleneck directly by restructuring the workflow around repeated matrix operations, which should significantly reduce the time spent on phenotype processing when you have thousands of traits.

Yuki: And from a population genetics viewpoint, if this method makes large-scale screening practical for complex trait panels, it opens up possibilities for studying how genetic architecture varies across different populations much faster than before.

Ines: That's the biological payoff I'm interested in—what kind of actual biological information do we recover when we run these analyses so rapidly?

Marcus: It recovers genotype-phenotype correlations and t-statistics all at once, giving you a consolidated view of where associations lie across all those traits simultaneously instead of doing them sequentially.

Yuki: That simultaneous view is really important because it allows for a much richer context when we're trying to prioritize specific traits within a species' genetic landscape.

Ines: So, while we’ve seen the mechanics, what are the real-world implications of having this kind of rapid screening tool available right now?

Marcus: The practical impact is definitely making high-throughput quantitative trait association much more accessible in settings like imaging genetics where you have a massive number of candidate traits and a slow pipeline just won't cut it.

Yuki: I see it as accelerating discovery in evolutionary studies; we can test hypotheses about complex trait inheritance with much greater statistical power by screening a wider range of traits concurrently.

Ines: It really puts the focus back on interpreting those results, rather than getting bogged down managing massive computational pipelines for every single analysis.

Department of Bioinformatics and Systems Medicine, McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston

cs.DC, cs.SE, q-bio.GN

Submitted: 2026-04-22

Updated: 2026-10-07

Comments: 5 pages, 2 figures

Code: https://github.com/ZhiGroup/TorchGWAS

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 77/100

The gist: TorchGWAS is a framework designed for high-throughput association testing across large panels of quantitative traits by leveraging GPU acceleration to overcome computational bottlenecks in

Key concepts

Batch Processing of Genotypes
Instead of running a separate analysis for each trait, TorchGWAS groups markers into batches. This allows the framework to perform repeated matrix operations efficiently across all traits simultaneously, which is crucial for scaling up to thousands of phenotypes.
Phenotype Residualization
This step cleans the phenotype data by removing effects from covariates using an orthonormal basis. It standardizes and residualizes the phenotype matrix before correlation calculations, ensuring that subsequent association tests are based on the underlying genetic signal rather than noise or confounding variables.
Genotype-Phenotype Correlation Matrix (R)
This matrix links every genotype to every phenotype. By computing this correlation, TorchGWAS creates a structure where researchers can easily calculate t-statistics for associations between specific markers and specific traits in a single vectorized operation.
Throughput Scaling
The framework scales well because the heavy preprocessing steps are done only once. This means the time taken to add more traits increases much slower than linearly with the number of phenotypes, making it highly practical for large-scale screening tasks.

Terminology

Summary

TorchGWAS is a framework designed for high-throughput association testing across large panels of quantitative traits by leveraging GPU acceleration to overcome computational bottlenecks in phenotype-rich screening workflows. The current public release provides stable Python and command-line workflows that allow researchers to reuse genotype matrices efficiently across thousands of phenotypes, making large-scale GWAS practical in settings like imaging genetics where many candidate traits must be evaluated rapidly.

How it works

TorchGWAS is implemented in Python and organized as a packaged command-line workflow based on the observation that the core workload can be reorganized as repeated matrix operations between batches of genotypes and a phenotype matrix. Instead of running an association pipeline for one phenotype at a time, TorchGWAS loads the entire phenotype panel, performs preprocessing once, and then reuses this processed matrix across the genome scan.

The association kernel involves several key steps:

  1. For an input phenotype matrix Y ∈ R!×P with P phenotypes over N samples, TorchGWAS first mean-centers each phenotype and removes covariate effects using an orthonormal basis Q spanning the covariate space. This residualization is performed as:

  2. Y' = (I−QQ')(Y−Ȳ) (1)

  3. The phenotype matrix is then standardized column-wise to unit variance, yielding Y'.

  4. Genotypes are processed in batches of M markers. For each batch, each marker vector is standardized across samples.

  5. TorchGWAS then computes the genotype-phenotype correlation matrix as R = G',Y / N (2), where R ∈ R(×P).

  6. These correlation coefficients are converted to t-statistics using T = R'3 !) +) (3).

Inputs and Preprocessing

The framework accepts several types of inputs, including:

(GENOTYPE)

(NumPy matrix)

(PLINK bed/bim/fam)

(BGEN + sample)

And for the phenotype data:

(Matrix or tabular file - Quantitative traits)

The preprocessing pipeline involves several critical steps to ensure efficiency and stability:

  1. Infer genotype sample order.

  2. Align phenotype/covariate tables by IID.

  3. Drop zero-variance columns.

  4. Compute covariate QR basis internally.

  5. Residualize and standardize phenotypes, resulting in outputs such as phenotype processed.npy, covariate q.npy, qc.json, prep.json.

Linear GWAS Output

The output of the linear GWAS process is structured to provide comprehensive association statistics simultaneously:

(results.tsv.gz)

(run.json)

(qc.json)

The results include:

  1. beta, se, t, -log10 P.

  2. Chunked genotype scan.

  3. Marker-wise correlation / t-statistic.

  4. Trait-specific P values.

This matrix formulation allows phenotype preprocessing and genotype traversal to be shared across large phenotype panels, which is what substantially improves throughput in phenotype-rich settings. The resulting statistics can then be used to rank associated loci, or incorporated into downstream phenotype-level prioritization.

Benchmark Results and Scaling

The performance of TorchGWAS was benchmarked against archived fastGWA results using a dataset of 8.9 million markers and approximately 23,000 samples. The comparison showed significant speedup: fastGWA required approximately 100 s per phenotype on an AMD EPYC 7763 64-core CPU, whereas TorchGWAS completed 2,048 phenotypes in about 940 s and 20,480 phenotypes in about 1108 s on a single NVIDIA A100 GPU. This corresponds to an approximately 300- to 1700-fold increase in phenotype throughput. The runtime increased much more slowly than linearly with the number of phenotypes because the preprocessing and genotype traversal steps are shared across the phenotype matrix, making it well-suited for settings where the main scaling challenge is the number of traits rather than the number of variants alone.

Scope and Limitations

The current version of TorchGWAS is intended for high-throughput quantitative-trait association under a linear model formulation. Its principal strength lies in throughput and workflow stability when many phenotypes share the same cohort and genotype matrix. However, two caveats are noted:

  1. the current release focuses on linear and multivariate quantitative-trait screening and does not yet provide the broader modelling options available in mature mixed-model toolkits.

  2. "large-scale end-to-end performance still depends on backend reader efficiency, storage bandwidth, and the execution environment.

Improvements for AI systems

Here are the specific improvements that could be made to AI systems by leveraging the TorchGWAS framework, along with what those improved systems could achieve:

  1. Improved Efficiency in High-Dimensional Phenotype Screening:

  2. Enhanced Scalability for Representation Learning and Imaging Genetics:

  3. Increased Throughput for Multi-Trait Discovery in Large Cohorts:

  4. Reduced Computational Bottlenecks in Deep Phenotyping Pipelines:

  5. Streamlined Workflow Integration with Genotype Data Formats (PLINK, BGEN, NumPy):


  1. Improved Efficiency in High-Dimensional Phenotype Screening: The system can now perform association testing across thousands of quantitative traits simultaneously by reusing the same preprocessed phenotype matrix for every genotype batch.

  2. Enhanced Scalability for Representation Learning and Imaging Genetics: The framework enables the efficient screening of genetic associations within learned representations (e.g., brain MRI features) where a single genotype matrix must be tested against a massive panel of derived quantitative traits, overcoming the computational barrier that limits current methods like tensorQTL.

  3. Increased Throughput for Multi-Trait Discovery in Large Cohorts: The system can rapidly generate thousands of trait-specific P-values from a single, large genotype scan (e.g., 8.9 million markers), significantly increasing the number of phenotypes evaluable per unit of time compared to existing tools that test traits sequentially.

  4. Reduced Computational Bottlenecks in Deep Phenotyping Pipelines: By amortizing phenotype preprocessing and genome scanning costs across all traits, the system drastically reduces redundant computation, making it feasible to screen large panels of quantitative traits derived from a single cohort (like UK Biobank deep phenotyping) efficiently on GPU hardware.

  5. Streamlined Workflow Integration with Genotype Data Formats (PLINK, BGEN, NumPy): The system provides native input support for multiple common genotype formats and NumPy arrays, allowing researchers to seamlessly integrate existing data pipelines directly into the high-throughput association testing workflow without extensive format conversion overhead.

Abstract

Imaging, molecular, and machine-learning workflows can generate thousands of quantitative phenotypes in a single cohort, creating substantial computational and output bottlenecks when testing traits individually. TorchGWAS is a GPU-accelerated framework that uses batched operations for high-throughput, covariate-adjusted linear association testing across large panels of quantitative phenotypes. Across 500,036 allele-harmonized tests, TorchGWAS t statistics agreed with PLINK 2.0. On an NVIDIA H100 80-GB GPU with a 48-core Intel Xeon Gold 6442Y host and measured disk read and write rates of 5.98 and 1.49 GB/s, respectively, median end-to-end times for 4.57 billion associations (8,931,083 variants by 512 phenotypes in 35,365 samples) were 28.46 s for BED, 29.13 s for hard-call PGEN, 51.48 s for BGEN, and 58.95 s for dosage PGEN, including writing 36.7 GB of binary summary statistics. TorchGWAS provides an efficient Python-based framework for parallel fixed-effect association screening at biobank scale.TorchGWAS is implemented in Python and distributed as a documented source repository at https://github.com/ZhiGroup/TorchGWAS.

Related papers