DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction
Dong Xu, Zhangfan Yang, Jiantao Wu, Zexuan Zhu, Jianqiang Li, Junkai Ji
Shenzhen University · EasternDawn · University of Nottingham Ningbo
q-bio.BM, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 19 pages, 2 figures, with supplementary material
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its
Terminology
Summary
Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context. The paper states: "A degrader molecule should not be modeled as having a single context-independent degradation label; its measured effect depends jointly on the molecular structure of the degrader, the target protein, the recruited E3 ligase, and the cellular context in which the assay is performed."
A key challenge is that although public databases contain thousands of structured molecule–target–E3 records, degradation measurements are available for only a small fraction of them. Existing supervised approaches therefore leave most recorded chemical–biological relationships unused.
The paper notes that a substantial portion of available data contains chemical and biological context but lacks a complete degradation label.
The authors introduce DegradeQuery, a context-aware prediction framework that converts these label-missing records into a pretraining signal.
The core innovation is counterfactual tuple pretraining, which contrasts recorded tuples with alternatives formed by replacing the target, the E3 ligase, or both, enabling the model to learn contextual associations without assigning activity pseudo-labels.
The paper emphasizes a conceptual shift: This design redefines the role of unlabeled PROTAC data. Rather than treating unlabeled records as incomplete supervised examples, DegradeQuery treats them as structured tuple observations.
The authors explicitly state: DegradeQuery does not generate pseudo-labels for unlabeled records and does not rely on teacher models, distillation, or ensembles. Unlabeled data enter the learning process solely through their observed tuple structure.
The labeled and unlabeled sets are defined as:
-
DL = (mi, ti, ei, ci, yi) for labeled records
-
DU = (mi, ti, ei, ci) for unlabeled records
where mi is the PROTAC molecule, ti is the target protein, ei is the E3 ligase, ci is an optional assay descriptor, and yi ∈ 0,1 is the binary degradation label.
The pretraining tuple set is constructed as: T = (mi, ti, ei): (mi, ti, ei, ci, yi) ∈ DL ∪ (mi, ti, ei): (mi, ti, ei, ci) ∈ DU. The paper clarifies: In this stage, a positive tuple means that the molecule–target–E3 context was recorded in the dataset; it does not mean that the tuple is degradation-active.
The model encodes molecule, target protein, and E3 ligase separately:
-
The molecule encoder fm
combines a three-layer residual message-passing network over 15-dimensional atom features with a 2,048-bit Morgan fingerprint (radius 2)
-
The shared protein encoder fp
maps 150 sequence features—amino-acid composition, sequence-length and unknown-residue features, and 128 hashed di-/tri-peptide counts—to 256 dimensions
The fusion module exposes both component identities and pairwise interactions
through concatenation of molecule, target, and E3 embeddings, their pairwise products, and the absolute target–E3 difference.
For each recorded tuple, counterfactual alternatives are constructed by replacing the target, the E3 ligase, or both. The counterfactual set is:
Ni = (mi, t̃i,r, ei), (mi, ti, ẽi,r), (mi, t̃i,r, ẽi,r) for r = 1 to R
The paper emphasizes: These counterfactuals are not assumed to be biologically impossible or degradation-inactive. They are sampler-defined alternatives that may include untested but viable molecule–target–E3 combinations.
The tuple scoring function is optimized with a contrastive ranking loss (noise-contrastive rather than supervised with true negative biological examples). A molecule-level consistency objective is also employed as an auxiliary regularizer, following graph and molecular contrastive learning approaches.
After pretraining, the model is fine-tuned on labeled degradation records using a class-weighted focal loss.
The paper notes: When γ = 0, this formulation reduces to weighted binary cross-entropy, so it subsumes the standard case as a special case of focal loss.
On the official PROTAC-8K benchmark split, DegradeQuery-CTP achieves an AUROC of 0.9065 and accuracy of 0.8500, outperforming all compared methods. Specifically, it improves over DegradeMaster-Semi (which uses pseudo-labeling) by 0.0134 accuracy and 0.0240 AUROC, and over DegradeQuery-Sup (supervised only) by 0.0367 accuracy and 0.0238 AUROC.
-
Five-seed paired evaluation:
Across five paired runs, every AUROC difference between CTP and Sup is positive; the mean paired gain is 0.0219 with a 95% confidence interval of [0.0136, 0.0303].
-
Unlabeled-only pretraining control: When every labeled row is removed from pretraining (using only 7,134 label-missing records), the model reaches
0.9007±0.0032 AUROC, improving over Sup by 0.0230 with a paired 95% confidence interval of [0.0173, 0.0287].
The difference from full CTP is only 0.0011, showingthat the label-missing records alone recover the full downstream improvement within the uncertainty of these runs.
-
ESM2-650M comparison: With frozen ESM2-650M protein representations, CTP improves AUROC from 0.8863 to 0.9032 compared to the matched supervised configuration, demonstrating that
ESM2 supplies stronger component representations while CTP continues to provide a useful pretraining signal for the full molecule–target–E3 predictor.
The paper compares tuple pretraining with molecule-only self-supervision: "Molecule-only SSL improves over labeled-only training, while tuple-only pretraining gives the higher AUROC and AUPRC of the two individual components. Combining the objectives gives the highest AUROC, indicating that the tuple objective contributes beyond generic molecular consistency."
A shortcut control shows that a molecule-ablated model using only target, E3, and assay-side features performs poorly, with AUROC near 0.55 and MCC near zero across three seeds. Thus, the gain is not explained by a target–E3 identity shortcut.
-
Target–E3 group holdout: Across ten independently constructed 20% group holdouts,
mean AUROC increases from 0.6769 to 0.7064, but the paired difference is heterogeneous: +0.0295 with a 95% confidence interval of [-0.0219, 0.0810], and five of ten splits have a positive AUROC difference.
The paper notes thisshows that CTP can help under biological-context shift while also identifying held-out context composition as an important source of uncertainty.
-
Scaffold holdout: CTP improves AUROC from 0.8681 to 0.8858 and MCC from 0.5919 to 0.6358,
indicating that it does not simply memorize frequent scaffolds.
-
Target-wise few-shot adaptation: Across K = 2, K = 4, and K = 8,
CTP improves every reported metric; the largest fixed-threshold gain appears at K = 2 in MCC, increasing from 0.1263 to 0.3521.
The paper acknowledges several limitations: "This study remains retrospective and benchmark-limited. Recorded tuples may encode database, publication, and medicinal-chemistry selection biases, and sampled counterfactuals may include untested but viable combinations. The authors also note that
A shuffled-pair control does not isolate the semantics of the observed target–E3 pairing as the unique source of the gain and that
repeated target–E3 holdouts reveal substantial context-dependent variation."
The paper concludes: "DegradeQuery should therefore be viewed as a data-driven context-aware predictor rather than a mechanism-aware model of ternary-complex geometry, ubiquitination, permeability, E3 expression, or cell-line-specific biology. Future work should combine harder biologically informed counterfactuals with harmonized external benchmarks and prospective validation."
The paper makes three main contributions:
-
We formulate unlabeled PROTAC records as structured molecule–target–E3 observations rather than incomplete examples awaiting pseudo-labels.
-
"We propose DegradeQuery, which employs counterfactual tuple pretraining to learn tuple-conditioned PROTAC representations without assigning degradation pseudo-labels, using teacher models, applying distillation, or ensembling models."
-
We isolate the contribution of label-missing records through five-seed paired evaluation, an unlabeled-only pretraining control, a molecule-versus-tuple component analysis, and an ESM2-650M comparison.
Improvements for AI systems
Improvements to AI Systems:
-
Context-Aware Pretraining from Unlabeled Structured Data: Replace pseudo-labeling or teacher-student approaches with counterfactual tuple pretraining that contrasts recorded (molecule, target, E3) tuples against alternatives formed by swapping components. This allows the model to learn contextual associations from label-missing records without fabricating supervision signals, improving data efficiency and avoiding confirmation bias from incorrect pseudo-labels.
-
Noise-Contrastive Ranking for Biological Context Learning: Implement a contrastive ranking loss over tuple structures (rather than binary classification) so the model learns which molecule–target–E3 combinations are plausible in the observed data distribution. This enables the AI to capture joint dependencies between chemical structure and biological context, improving generalization to untested combinations.
-
Component-Aware Fusion Architecture: Use separate encoders for molecule, target, and E3, then fuse via concatenation plus pairwise interactions and absolute differences. This forces the model to explicitly represent both individual component identities and their interactions, avoiding shortcut learning from any single feature (e.g., target identity alone) and improving robustness to context shifts.
-
Auxiliary Molecule-Level Consistency Regularization: Add a self-supervised consistency objective on molecular representations alongside the tuple objective. This provides complementary signal that improves performance beyond either objective alone, as shown by the component ablation, and helps the model learn more transferable molecular features.
-
Class-Weighted Focal Loss for Fine-Tuning: Use focal loss with class weighting during supervised fine-tuning to handle class imbalance in degradation labels. This subsumes standard binary cross-entropy as a special case (γ=0) and improves calibration on rare positive/negative degradation outcomes.
-
Unlabeled-Only Pretraining as a Data-Efficient Strategy: The finding that pretraining on only label-missing records recovers nearly the full improvement (AUROC 0.9007 vs 0.9065 for full CTP) suggests AI systems can be pretrained entirely on unlabeled structured data before any labeled fine-tuning. This enables deployment in domains where labeled data is scarce but structured records are abundant.
-
Few-Shot Adaptation with Pretrained Contextual Representations: The significant gains in low-data regimes (e.g., MCC improvement from 0.1263 to 0.3521 at K=2) indicate that counterfactual tuple pretraining produces representations that adapt rapidly to new targets or contexts with very few labeled examples. This is critical for personalized medicine or rare disease applications where per-target data is limited.
-
Scaffold-Agnostic Generalization: The improvement on scaffold holdout (AUROC 0.8681→0.8858) shows the pretraining prevents overfitting to frequent chemical scaffolds, enabling the AI to predict degradation for novel molecular structures—essential for virtual screening of new PROTAC candidates.
-
Explicit Uncertainty Quantification for Context Shift: The heterogeneous results on target–E3 group holdouts (mean gain +0.0295 AUROC but wide confidence interval) suggest the AI should be trained to output calibrated uncertainty estimates, flagging predictions where the biological context is novel or underrepresented, rather than giving overconfident point estimates.
-
Bias-Aware Training with Counterfactual Sampling: Since recorded tuples may contain database or publication biases, the counterfactual sampling strategy should be explicitly designed to diversify the negative space (e.g., sampling from chemically similar but untested molecules, or biologically plausible but unrecorded target–E3 pairs). This reduces the risk of learning dataset-specific artifacts rather than true biological relationships.
What the Improved AI System Can Do:
-
Predict PROTAC degradation efficacy for novel molecule–target–E3 combinations with higher accuracy (AUROC 0.9065 on PROTAC-8K) than supervised-only or pseudo-labeling approaches.
-
Learn effectively from unlabeled structured data (thousands of records without degradation labels) without requiring manual annotation or pseudo-label generation.
-
Adapt to new protein targets or E3 ligases with only 2–8 labeled examples, achieving substantial gains in MCC (e.g., 0.1263→0.3521 at K=2) for few-shot drug discovery scenarios.
-
Generalize to novel chemical scaffolds and biological contexts, avoiding memorization of frequent training combinations.
-
Provide context-aware predictions that jointly consider molecular structure, target protein, and E3 ligase identity, rather than treating degradation as a molecule-only property.
-
Operate in data-scarce regimes by pretraining entirely on unlabeled structured records, then fine-tuning on minimal labeled data—applicable to other domains with similar tuple-structured data (e.g., drug–target–cell line interactions, protein–protein interaction modulators).
Sources
- Representation Learning with Contrastive Predictive Coding
- Context-enriched molecule representations improve few-shot drug discovery
Related papers
- Speak to a Protein: An Interactive Multimodal Co-Scientist
- Learning Topological Representations of Protein Structure and Dynamics
- p2smi: A Python Toolkit for Peptide FASTA-to-SMILES Conversion and Molecular Property Analysis
- UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions
- Co-folding with a Soup of Representations
- Non-Markovain Quantum State Diffusion for the Tunneling in SARS-COVID-19 virus