Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

arXiv:2608.11444 · q-bio.QM, cs.LG · Submitted 2026-08-17 · Read on arXiv

Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor, Thomas Brettin, Rick Stevens

Argonne National Laboratory · University of Chicago

q-bio.QM, cs.LG

Submitted: 2026-08-17

Updated: 2026-08-19

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper presents a substantial expansion of the IMPROVE benchmark dataset for drug response prediction (DRP) models, integrating large-scale pharmacogenomic data primarily from PharmacoDB along

Terminology

Summary

The paper presents a substantial expansion of the IMPROVE benchmark dataset for drug response prediction (DRP) models, integrating large-scale pharmacogenomic data primarily from PharmacoDB along with additional smaller data sources. The expanded resource includes millions of drug–response measurements, broader multi-omics coverage, and a major increase in chemical diversity, adding more than 50,000 compounds. All data were curated and standardized using consistent data preprocessing, molecular featurization, and dose–response modeling procedures to produce a unified, AI-ready dataset suitable for drug response modeling.

The dataset integrates data from multiple sources: CCLE, CTRPv2, GDSCv1 and GDSCv2, gCSI, NCI-60, PRISM, FIMM, and several patient-derived organoid (PDO) studies (Lee et al., Narasimhan et al., Zhu et al., Van de Wetering et al.). The total dataset includes 1,362 unique cancers, 53,949 drugs, and 5,455,444 drug–response experiments. The NCI60 study contributes the majority of the data, with 53,323 drugs and 4,267,356 experiments.

Drug response data were processed using a unified workflow. Response values were truncated to an upper bound of 200 percent viability prior to curve fitting. Drug response was summarized by fitting a four-parameter logistic (4PL) dose–response model to normalized viability measurements as a function of log-transformed drug concentration. The model is defined as f(x) = E∞ + (E0 − E∞)/(1 + 10(x−EC50)·HS), where x denotes log10-transformed molar concentration, E0 and E∞ represent asymptotic response at zero and infinite dose, EC50 is the half-maximal response concentration, and HS is the hill slope. Two fitting regimes were employed: a strict regime enforcing canonical decreasing response with bounds E∞ ∈ [0, 1], EC50 ∈ [−19, 5], HS ∈ [−5, 5], E0 = 1, and a flexible regime allowing inverted dose–response relationships with relaxed bounds E∞ ∈ [0, ∞), EC50 ∈ [−19, 5], HS ∈ [0, 5], E0 ∈ [0, ∞). Selection between fits was based on R2, with the inverted fit retained only if it improved R2. Area under the dose response curve (AUC) was computed under both regimes, with extreme AUC values capped at 1.0.

Chemical data were standardized by cleaning SMILES strings, removing salts and counterions using RDKit's SaltRemover, generating canonical SMILES, and collapsing tautomeric variants. Two molecular feature representations were derived: physicochemical descriptors computed using the Mordred library (with 3D descriptors disabled, zero-filling undefined values) and extended-connectivity fingerprints (ECFP4) with radius 2 and fixed length of 512 bits.

The multi-omics data include six modalities: gene expression, copy number variation, somatic mutation, DNA methylation, protein expression (RPPA), and miRNA expression. Data were sourced from the Dependency Map (DepMap) portal of CCLE version 22Q2, with gene expression supplemented from PharmacoDB for cell lines not in DepMap. Gene expression profiles were converted to log2(TPM + 1) and quantile normalized. Copy number variation values were represented as log2(CN ratio + 1) with discretized versions using five states. Mutation data include both a gene-level count matrix and a long-format table with individual variant records. DNA methylation data were generated using reduced-representation bisulfite sequencing (RRBS) with features corresponding to gene transcription start sites. RPPA measures protein abundance and post-translational modification for a targeted panel of proteins. miRNA expression quantifies abundance of individual miRNAs. Gene identifiers were mapped using NCBI, and cell line identifiers were mapped to Cellosaurus accession IDs.

Two DRP models were evaluated: UNO, which uses separate fully connected subnetworks to encode cell line and drug features (drug features represented using Mordred physicochemical descriptors, cell line features using gene expression), and GraphDRP, which represents drugs as molecular graphs and uses graph convolutional networks, with gene expression as cell line input.

Model performance was evaluated using 10-fold cross-validation with evaluation folds defined exclusively using experiments from GDSC, CCLE, and CTRPv2. Drug-blind test sets contained experiments involving held-out drugs, cancer-blind test sets contained experiments involving held-out cell lines, disjoint test sets contained experiments where both drug and cell line were unseen, and mixed test sets contained experiments where drug and cell line were seen but their pairing was not. NCI60, PRISM, and FIMM were used only for training. Training was performed on an 8-GPU cluster with batch size 512, mean squared error loss, and early stopping with patience of 20 epochs.

Results showed substantial improvements in drug-blind and disjoint settings when trained on the expanded dataset. For UNO, mean R2 in drug-blind improved from 0.03 to 0.22, and in disjoint from-0.10 to 0.20. For GraphDRP, drug-blind R2 improved from-0.11 to 0.12, and disjoint from-0.26 to 0.08. Mean PCC for UNO improved from 0.39 to 0.51 in drug-blind and from 0.34 to 0.49 in disjoint; for GraphDRP, PCC improved from 0.21 to 0.45 in drug-blind and from 0.14 to 0.42 in disjoint. However, performance slightly decreased in mixed and cancer-blind settings. For UNO, mixed R2 decreased from 0.71 to 0.67, and cancer-blind from 0.60 to 0.58. For GraphDRP, mixed R2 decreased from 0.75 to 0.66, and cancer-blind from 0.60 to 0.57. The paper suggests the slight performance drop in mixed and cancer-blind settings may be due to the NCI60 dataset dominating the expanded dataset and using a different viability assay than the test datasets (GDSCv2, CCLE, CTRPv2).

The paper concludes that the expanded dataset provides a stronger foundation for developing DRP models that generalize to novel compounds, which is critical for drug discovery applications. Limitations include increased computational resource requirements, residual noise from heterogeneous experimental sources, and the fact that additional omics modalities were not fully leveraged in the current analysis. The paper suggests future work could incorporate multi-omics integration and explore pretraining strategies and foundation-model approaches.

Improvements for AI systems

Improvements to AI systems:

  1. Enhanced generalization to unseen drugs: By training on the expanded dataset (5.4M experiments, 53,949 drugs), the AI system can now predict drug responses for novel compounds with significantly higher accuracy (UNO drug-blind R2 improved from 0.03 to 0.22; PCC from 0.39 to 0.51). The improved system can screen large chemical libraries for candidate therapeutics without requiring prior experimental data on those specific drugs.

  2. Robust handling of heterogeneous experimental data: The unified preprocessing pipeline (4PL curve fitting with strict/flexible regimes, R2-based selection, capped AUC, standardized SMILES cleaning, and consistent multi-omics normalization) enables the AI to learn from noisy, multi-source pharmacogenomic data. The improved system can automatically reconcile dose–response measurements from different assays (e.g., NCI60 vs. GDSC) into a coherent representation, reducing batch effects and assay-specific biases.

  3. Multi-omics feature integration readiness: The dataset provides six aligned omics modalities (gene expression, CNV, mutation, methylation, RPPA, miRNA) with standardized identifiers (Cellosaurus, NCBI). An improved AI system can be trained to fuse these modalities—e.g., using a multi-modal encoder that jointly processes transcriptomic, proteomic, and epigenetic inputs—to predict drug response with higher accuracy than single-omics models, especially for cancer types where gene expression alone is insufficient.

  4. Chemical representation flexibility: The provision of both Mordred physicochemical descriptors and ECFP4 fingerprints (512-bit) allows the AI to use complementary drug representations. An improved system can employ a dual-encoder architecture that learns from both descriptor-based and graph-based features, or use contrastive learning to align these representations, improving robustness across diverse chemical scaffolds.

  5. Domain-adaptive training strategies: The observed performance drop in mixed/cancer-blind settings (due to NCI60 dominance) can be mitigated by an improved AI system that uses domain adaptation or reweighting techniques—e.g., adversarial training to align feature distributions between NCI60 and GDSC/CCLE/CTRPv2, or meta-learning to explicitly penalize overfitting to the dominant source. This would preserve gains in drug-blind/disjoint settings while maintaining or improving cancer-blind and mixed performance.

  6. Foundation-model pretraining on curated dose–response data: The large, standardized dataset enables pretraining of a foundation model for drug–cell interactions. The improved system can be pretrained on the full 5.4M experiments using self-supervised objectives (e.g., masked drug/cell feature reconstruction, dose–response curve prediction) and then fine-tuned on smaller, specialized datasets (e.g., patient-derived organoids) for personalized medicine, requiring far fewer labeled examples.

  7. Improved uncertainty quantification: The dual-fitting regime (strict vs. flexible) and capped AUC values provide a natural basis for modeling aleatoric uncertainty. An improved AI system can output predictive distributions (e.g., via heteroscedastic regression or Monte Carlo dropout) that reflect confidence in dose–response predictions, enabling safer decision-making in drug discovery—e.g., flagging low-confidence predictions for experimental validation.

  8. Scalable training with early stopping and cross-validation: The demonstrated 10-fold CV protocol (with drug-blind, cancer-blind, disjoint, and mixed splits) can be embedded into an automated training pipeline. The improved AI system can automatically select the best model checkpoint based on drug-blind/disjoint performance (not just mixed), ensuring deployment-ready models that prioritize generalization to novel compounds and unseen combinations.

What the improved AI system can do:

  • Virtual screening of millions of compounds against thousands of cancer cell lines and organoids, prioritizing novel drug candidates with high predicted efficacy and low toxicity, even for chemical structures never seen in training.

  • Personalized therapy recommendation by integrating a patient’s tumor omics profile (expression, mutation, methylation, etc.) with drug features to predict response, with confidence intervals, for both approved drugs and investigational agents.

  • Cross-assay harmonization—automatically correcting for platform-specific biases (e.g., NCI60 vs. PRISM) when pooling data from multiple pharmacogenomic studies, enabling meta-analyses and larger training cohorts.

  • Rapid adaptation to new cancer types via transfer learning from the pretrained foundation model, requiring only a few dozen patient-derived organoid samples to achieve clinically useful accuracy.

  • Mechanistic insight generation by analyzing which omics modalities and chemical features contribute most to predictions (e.g., via attention or SHAP), helping researchers identify biomarkers of drug sensitivity and resistance.

Abstract

Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. However, their predictive performance is often constrained by limited dataset scale and insufficient coverages of cancer and chemical spaces. In addition, inconsistent benchmarking practices hinder reliable comparison across models. Standardized frameworks, such as the Innovative Methodologies and New Data for Predictive Oncology Model Evaluation (IMPROVE) project, provide unified data schemas and evaluation protocols for consistent benchmarking, but improving model generalizability requires larger and more diverse training data. In this work, we substantially expand the IMPROVE benchmark through large-scale integration of pharmacogenomic data, primarily from PharmacoDB, together with additional smaller data sources. The expanded resource includes millions of drug response measurements, broader multi-omics coverage, and a major increase in chemical diversity, adding more than 50,000 compounds. To evaluate the impact of the new dataset compared to the original IMPROVE benchmark dataset, we trained DRP models using the two datasets and assess their prediction performance using a common test set and several evaluation strategies, including drug-blind, cancer-blind, and disjoint data splits. While cancer-blind performance remained comparable to the original benchmark, models trained on the expanded dataset showed consistent improvements in drug-blind and disjoint settings, indicating enhanced generalization to previously unseen compounds. These results position the expanded dataset as a community resource that provides a richer foundation for developing DRP models intended to aid in the discovery of novel anticancer drugs.

Related papers