Learning to Select Source Domains: Proxy-Rewarded Policy Optimization for Molecular OOD Generalization

arXiv:2605.13932 · cs.LG · Submitted 2026-05-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Learning to Select Source Domains".

Tom: Robust prediction of molecular properties under extreme out-of-distribution (OOD) scenarios is a pivotal bottleneck in AI-driven drug discovery,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, we've just been diving deep into how this paper tackles out-of-distribution problems in molecular AI by intelligently picking which data sources to use, and now we're wrapping up with a look at what this whole piece means for the future.

Jane: Exactly. We’ve seen how they use policy optimization to guide knowledge transfer, and now it’s time for us to really unpack the title and who came up with this important work.

Lu: I think focusing on "Learning to Select Source Domains" is crucial because it moves the focus away from just making bigger models, toward designing smarter data pipelines, which is a wild direction for molecular modeling.

Meng: From an engineering standpoint, understanding who wrote this helps me gauge the rigor; it’s important to know what kind of team developed these sophisticated selection policies that we might try to build ourselves someday.

Lalam: The authors themselves are presenting a framework that links target awareness directly into the learning process, which suggests a cultural shift where data curation becomes an active, learned component of the AI system itself.

Tom: And what they're saying is pretty straightforward: this research is about using targeted selection to stop models from failing when they hit new chemical structures completely out of their training range.

Jane: That’s the essence of it—it’s about giving the AI a smarter way to choose its reading material so it doesn't get confused by novel inputs.

Lu: It really opens up possibilities for exploring chemical spaces we couldn't reach before because the model is now guided to pull from just the right neighborhood of knowledge.

Meng: I wonder how this dynamic selection policy translates into a stable, deployable system in a real-world drug discovery pipeline, given all those complex components like GRPO and dual-scale alignment.

Lalam: If this works as described, it could fundamentally change how we approach AI development across different domains, suggesting that intelligent knowledge orchestration is the next major evolution in making AI truly useful.

Conclusion: Tom: So, we've just finished looking at all the technical details of this paper about intelligently selecting data sources for molecular AI, and now we need to talk about what this whole title actually means for us.

Jane: You're right, Tom; essentially, the paper is showing us a way to make models more resilient when they encounter molecular structures they haven't seen before in their training set.

Lu: I think focusing on "Learning to Select Source Domains" is crucial because it moves the focus away from just making bigger models, toward designing smarter data pipelines, which is a wild direction for molecular modeling.

Meng: From an engineering standpoint, understanding who wrote this helps me gauge the rigor; it’s important to know what kind of team developed these sophisticated selection policies that we might try to build ourselves someday.

Lalam: The authors themselves are presenting a framework that links target awareness directly into the learning process, which suggests a cultural shift where data curation becomes an active, learned component of the AI system itself.

Tom: And what they're saying is pretty straightforward: this research is about using targeted selection to stop models from failing when they hit new chemical structures completely out of their training range.

Jane: That’s the essence of it—it’s about giving the AI a smarter way to choose its reading material so it doesn't get confused by novel inputs.

Lu: It really opens up possibilities for exploring chemical spaces we couldn't reach before because the model is now guided to pull from just the right neighborhood of knowledge.

Meng: I wonder how this dynamic selection policy translates into a stable, deployable system in a real-world drug discovery pipeline, given all those complex components like GRPO and dual-scale alignment.

Lalam: If this works as described, it could fundamentally change how we approach AI development across different domains, suggesting that intelligent knowledge orchestration is the next major evolution in making AI truly useful.

Tom: It seems like they've validated this concept through extensive experiments, demonstrating up to an eleven point two percent reduction in mean absolute error with an average relative improvement of six point two percent across all tasks.

Jane: That performance gain across different backbone architectures suggests the method has a level of universality that is quite promising for real-world applications.

Lu: The authors, Zhuohao Lin, Kun Li, Jiameng Chen, and Wenbin Hu, are proving that source composition matters more than just picking the largest model because they show how to overcome negative transfer through a target-aware selection policy and dual-scale decoupled domain adaptation.

Meng: Practically speaking, for an engineer like me, this means we move away from static pre-processing steps toward a dynamic system that chooses the best data mix in real time.

Lalam: If we can build systems where the AI dynamically decides which knowledge to draw from based on what it needs to predict, it could foster an entirely new level of adaptability in how these models learn and apply their knowledge across different chemical spaces.

Tom: So to wrap up on the implications of this work, it suggests that controlling the quality and relevance of training data sources is a primary driver for molecular AI success.

Jane: In simpler terms, they're showing us how to stop models from getting confused when they see something completely new by making them smarter about which old pieces of knowledge to pull from.

Lu: The dual-scale decoupling is key because it ensures that this selection policy is supported by an adaptation process that maintains both the high-level structure and the low-level chemistry accurately.

Meng: From a practical implementation standpoint, ensuring that the dynamic weight controller effectively balances those regression losses against the alignment scales will be a significant engineering challenge we'll have to tackle.

Lalam: I think the cultural impact is huge; it shows that AI development isn't just about scaling up models, but about designing intelligent control mechanisms that manage complexity and uncertainty gracefully.

Tom: That seems to be the core message: by optimizing the selection process itself, we can handle those extreme structural shifts much better than before.

Jane: And this leads us perfectly into what this means for drug discovery, which is our next big topic.

School of Computer Science, Wuhan University · Department of Data Science and Artificial Intelligence, Monash University · College of Computer Science and Technology, Zhejiang University · School of Life Sciences and Technology, Tongji University

cs.LG

Submitted: 2026-05-13

Updated: 2026-09-28

Importance score: 78/100

The gist: Robust prediction of molecular properties under extreme out-of-distribution (OOD) scenarios is a pivotal bottleneck in AI-driven drug discovery, and this work addresses it by proposing a framework

Key concepts

Scaffold Split Benchmark (SCOPE-BENCH)
A new testing standard designed specifically for molecular prediction that prevents models from cheating by using subtle structural overlaps between datasets. It enforces strict separation based on chemical descriptors, revealing how vulnerable current models are when faced with truly novel structures.
Policy Optimization for Multi-source Adaptation (POMA)
A framework that treats knowledge transfer as a decision-making process. Instead of using fixed methods, POMA learns a policy to dynamically select which source domains to use, creating an integrated pipeline for retrieving, composing, and adapting models.
Dual-Scale Decoupled Domain Adaptation
A technique used during training that aligns the model's features at two different levels: the macroscopic whole-molecule level and the microscopic pharmacophore fragment level. This ensures that both large structural features and fine chemical details are accurately transferred from source to target data.

Terminology

Summary

Robust prediction of molecular properties under extreme out-of-distribution (OOD) scenarios is a pivotal bottleneck in AI-driven drug discovery, and this work addresses it by proposing a framework that intelligently selects optimal source domains for knowledge transfer. The core finding is that target-aware source selection, implemented via policy optimization guided by proxy rewards, significantly mitigates negative transfer and catastrophic degradation observed in state-of-the-art models when faced with extreme structural shifts.

The gist

Policy optimization for multisource adaptation (POMA) formulates knowledge transfer as an integrated, policy-driven retrieve–compose–adapt pipeline that overcomes negative transfer through a target-aware selection policy and dual-scale decoupled domain adaptation.

Scaffold Split Benchmark (SCOPE-BENCH)

The paper introduces the scaffold-cluster out-of-distribution performance evaluation benchmark (SCOPEBENCH) to resolve the issue of microscopic semantic overlap in conventional benchmarks. This benchmark enforces strict metric separation based on explicit physicochemical descriptor clustering to completely preclude hidden structural interpolation, contrasting with conventional scaffold splitting protocols that fail to obstruct microscopic semantic overlap. The evaluation shows that prediction errors of state-of-the-art 3D molecular models surge by up to 8.0× on SCOPEBENCH with a mean of 5.9× when compared to standard splits, exposing their fundamental OOD vulnerability.

Policy Optimization for Multi-source Adaptation (POMA)

POMA is a framework that formulates knowledge transfer as an integrated, policy-driven retrieve–compose–adapt pipeline. This pipeline introduces two core innovations:

  1. Revolutionizing retrieval and composition by learning a combinatorial selection policy via Group Relative Policy Optimization (GRPO). Unlike static graph kernels, the policy dynamically explores an exponentially large combinatorial space to compose the most synergistic source subset without requiring a fragile value network.

  2. Innovating domain adaptation by replacing conventional global alignment with a dual-scale decoupled architecture that aligns macroscopic whole-molecule topologies and microscopic pharmacophore fragments independently, ensuring that the selection policy is supported by a robust adaptation process that preserves both structural and chemical precision.

Virtual Target and Source Construction

The framework constructs proxy targets for reward estimation under unlabeled target conditions to establish an effective gradient feedback path. This involves:

  1. Calculating the Morgan fingerprint cosine similarity between a candidate scaffold and real target scaffolds to define similarity, given by Equation (2).

  2. Defining a joint HubScore to comprehensively evaluate the suitability of a candidate scaffold as a proxy, ensuring it maintains broad isomorphic connections with multiple real targets: HubScore(s) = max ti∈Treal Sim(s, ti) + λ X ti∈Treal I(Sim(s, ti) > τsim).

  3. Ranking candidate source domains using the Srank metric: Srank(s∗, cj) = kW L(s∗, cj) × ln(1 + Dcj), where kW L is the Weisfeiler–Lehman graph kernel similarity and Dcj is the sample size of domain cj.

Dual-Scale Decoupled Domain Adaptation

To minimize distribution discrepancy while preserving fine-grained chemical semantics, POMA implements a dual-scale decoupled alignment strategy. This involves constructing two parallel feature paths: a macroscopic path yielding whole-molecule features (hmol) and a microscopic path yielding pharmacophore features (hsub) via BRICS retrosynthetic cleavage. The unified multi-source alignment objective is defined as: LDA = wmolX K k=1 γk ·∥Cmol,k s − Cmol t ∥2 F 4d 2 + wsubX K k=1 γk ·∥Csub,k s − Csub t ∥2 F 4d 2 (Equation (5)). The total training objective combines supervised regression with alignment: Ltotal = wreg Lreg + LDA (Equation (6). The training proceeds in two phases, using a dynamic weight controller to adaptively balance the regression loss and the alignment scales.

Source Domain Combinatorial Decision via GRPO

The decision network uses a state representation vector xj ∈ R dstate constructed from fingerprint features, candidate similarity scores, and domain sample size information (Equation (7)). A policy network outputs a Bernoulli selection probability pj for each candidate scaffold. The action vector is generated through independent Bernoulli sampling based on these probabilities. The GRPO optimization objective is designed to select the optimal source subset by maximizing the intra-group normalized advantage: Aˆ i = Ri − µRV σRV + ϵs (Equation (9). This process allows the policy to navigate the exponentially large combinatorial space, leading to a substantial performance boost on challenging targets such as Scaffold 7 and Scaffold 15.

Overall Results

Extensive experiments across three backbone architectures demonstrate up to an 11.2% reduction in mean absolute error with an average relative improvement of 6.

Improvements for AI systems

As a fastidious researcher, I have analyzed the core contributions of this work: SCOPE-BENCH (a rigorous OOD benchmark) and POMA (a policy-driven, dual-scale adaptation framework).

Here are the specific improvements and capabilities that can be achieved by integrating these methods into existing AI systems:


The proposed framework offers two major system upgrades: a better test mechanism and a more robust learning mechanism.

  1. Enhance Out-of-Distribution (OOD) Robustness in Molecular Property Prediction:

  2. Improve Knowledge Transfer Efficiency via Target-Aware Source Selection:

  3. Enable Cross-Architecture Generalization for Drug Discovery Models:

The improved AI system can perform the following specific tasks:

  1. Predict molecular properties (HOMO, LUMO, HOMO-LUMO Gap) with significantly reduced error (up to 6.2% relative improvement) when encountering novel chemical scaffolds or structural shifts that are far outside the training distribution, something standard models fail at.

  2. Automatically determine and select the optimal subset of source data for fine-tuning a model based on its current prediction task and target structure, effectively eliminating negative transfer caused by aligning heterogeneous source libraries.

  3. Achieve superior performance across different 3D molecular architectures (ViSNet, ETNN, GotenNet) without requiring architecture-specific tuning or re-training of the core feature extraction layers.

  4. Develop a policy-guided orchestration system where the model doesn't just adapt blindly; it actively perceives the target domain and intelligently navigates a combinatorial space of source domains to find synergistic knowledge transfer pathways (via GRPO).

This capability transforms molecular AI from a brittle, distribution-dependent tool into a resilient system capable of reliable extrapolation in real-world drug discovery scenarios.

Sources

Related papers