A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization
Kelvin P. Idanwekhai, Enes Kelestemur, Benjamin Strickland, Matthew Hart, Steini Davidsson, Angelos Angelopoulos, Ron Alterovitz, Marcello DeLuca, Alexander Tropsha
University of North Carolina at Chapel Hill · Eshelman School of Pharmacy at UNC · Department of Applied Physical Sciences, UNC · Department of Computer Science, UNC
cs.AI, cs.LG, q-bio.QM
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 22 pages, 9 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 48/100
The gist: SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration) is an open-source framework that employs natural-language orchestration to guide chemical structure optimization.
Terminology
Summary
SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration) is an open-source framework that employs natural-language orchestration to guide chemical structure optimization. It uses an LLM to interpret user-defined goals and route tasks, while specialized tools perform reaction-templated analog enumeration, physicochemical and ADMET property prediction, structure-based affinity scoring, and Bayesian optimization. The resulting workflow is a computational twin of the analytical and prioritization stages of the design–make–test–analyze cycle, providing provenance of each numerical output. Across single, and multi-objective optimization studies, SABLE enriches candidate sets for user-defined computational objectives while evaluating only a subset of the enumerated search space. Its modular architecture allows tools and characterization backends to be replaced by editing a simple config file, without modifying operational logic. SABLE provides an extensible decision-support framework for prioritizing synthetically constrained analogs in early-stage drug discovery.
The framework is designed to address the multi-objective nature of hit-to-lead optimization, which requires balancing potency, selectivity, physicochemical properties, pharmacokinetics, safety, and synthetic tractability. A key architectural design in SABLE is the distribution of the respective contributions between LLMs, QSAR, and optimization tools. The LLM does not optimize molecules but translates human intent into machine-readable optimization configurations; resolving molecule names to canonical SMILES, preparing node-routing configurations, and surfacing results back to the user. It is never used to propose analogs, score affinity, or reason about ADMET liabilities. Each of those decisions is delegated to a purpose-built component with validated statistical or physical grounding: HEALER for synthetically constrained enumeration, Bayesian Backend (BayBE) for BO, Boltz-2 for structure-based affinity, and dedicated QSAR predictors for ADMET endpoints. This responsibility split is what distinguishes SABLE from other systems like ChemCrow or Coscientist that place the LLM at the center of scientific reasoning. This design is informed by evidence that LLMs can hallucinate chemically invalid structures, misreport physicochemical properties, and are biased toward molecular patterns overrepresented in their training corpus.
In single-objective optimization experiments targeting human calcium/calmodulin-dependent protein kinase kinase 2 (CAMKK2; UniProt Q96RR4), the best-observed predicted logIC50 improved by 1.03 log units relative to the seed, while the cumulative distribution of the 20 best-observed candidates plateaued after approximately five iterations. Evaluated compounds were distributed across the projected analog space, although most of the enumerated library remained unevaluated at campaign termination. In dual-objective optimization jointly optimizing Boltz-2-predicted affinity and a quantitative estimate of drug-likeness (QED), the evaluated candidates produced a set of nondominated solutions spanning different compromises between lower predicted logIC50 and higher QED, with the resulting Pareto frontier providing a set of alternative candidates rather than a single optimum. In multi-objective optimization considering binding affinity, QED, and CNS Activity, early iterations emphasized exploration to establish an initial Pareto frontier, while later iterations concentrated on regions predicted to yield balanced improvements across several properties. The resulting enriched libraries contained molecules that simultaneously satisfied multiple drug-like criteria.
The efficiency of exploring the chemical space was benchmarked on one-million-compound libraries using QED as the objective. The optimizer recovered the global maximum after three iterations in the similarity-sampled library and seven iterations in the diversity-sampled library. The efficiency gain becomes even more pronounced as the number of objectives increases, since multi-objective optimization rapidly becomes intractable under exhaustive screening. Quantitatively, across the three campaign classes, SABLE evaluated only a fraction of the enumerated library before convergence, translating to a reduction in Boltz-2 inference calls relative to exhaustive screening. The relative savings in oracle or experimentation calls grow approximately linearly with library size in the regime of 10 cubed to 10 6.
In a retrospective medicinal chemistry setting using the METTL3 inhibitor campaign reported by Dolbois et al., SABLE was supplied with the starting compound from the paper, specified METTL3 as the target using UniProt ID Q86U44, and asked to minimize predicted binding affinity. The starting compound was predicted by Boltz-2 to have an IC50 of 9.77 µM, which is consistent with the reported low-micromolar experimental activity of the starting compound (7 µM). For the published optimized compound UZH2, Boltz-2 predicted an IC50 of 31.6 nM, whereas the experimental TR-FRET IC50 was 5 nM. Without using any experimental feedback from the METTL3 campaign, SABLE identified a top-ranked analog with a predicted concentration of 102.33 nM, representing approximately a 95-fold decrease in predicted concentration from a single optimization run. The candidate was not experimentally tested, so the result is interpreted as a prospective prioritization rather than experimental validation. SABLE did not rediscover the exact published lead compound because the HEALER library used in this run did not include the reaction template required to construct that final triazaspiro[5.5]undecan-2-one series.
In a validation case study on Boltz-2 tested targets, four targets were selected from the Boltz-2 affinity-validation set that span distinct target classes: beta-secretase 1 (P56817, protease), carbonic anhydrase XII (O43570, lyase), ABL1 (P00519, kinase), and sphingosine-1-phosphate receptor 1 (P21453, GPCR). For each assay, the censored-low-activity compound in the experimental series was used as the seed molecule, while the best ChEMBL compound in the same assay provided a reference point. Across all four targets, SABLE identified analogs with improved Boltz-predicted affinity relative to the starting compound. The largest predicted shift was observed for P56817, where the seed was predicted at 58.88 µM and the best SABLE analog at 370 nM. O43570 and P21453 also showed substantial predicted shifts, from 400 nM to 19.95 nM and from 1.38 µM to 74.44 nM, respectively. P00519 began from a seed that Boltz already scored as considerably stronger than its ChEMBL assay value, but SABLE still improved the predicted value from 97.72 nM to 52.97 nM. The final predicted values are also in the same broad range as the experimental activities of the best ChEMBL reference compounds for P56817, O43570, and P21453.
SABLE operationalizes the architecture as a stateful workflow implemented using LangGraph. Each campaign is initialized with the seed molecule, protein target, optimization objectives, iteration limit, and batch size extracted from the user request. A shared state records the enumerated candidate library, previously evaluated molecules, returned characterization values, and optimization history. During each iteration, the Bayesian optimization node recommends a batch of unevaluated candidates, which are routed to the characterization tools associated with the specified objectives. The returned values are appended to the campaign history and used to update the Gaussian-process surrogate before the next batch is selected. Conditional routing continues this cycle until the iteration limit is reached, the predefined convergence criterion is satisfied, or the enumerated search space is exhausted. LangGraph checkpointing preserves intermediate workflow states, enabling interrupted campaigns to resume without repeating completed characterization steps.
The argument extraction node employs a hybrid extraction mechanism that combines rule-based regex parsing with LLM-assisted interpretation. The LLM component can handle ambiguous or complex prompts, resolving molecule names to structures and inferring implicit optimization objectives. From a prompt, SABLE can extract the starting molecule (SMILES), protein target (UniProt ID), optimization target (binding affinity, minimize), enumeration size, iteration count, and batch size. SABLE automatically detects the number and types of objectives present in the user prompt during the argument extraction state. After argument extraction, SABLE runs a name-to-entity
step to normalize user-specified molecular and protein identifiers into machine-readable structures, using a hybrid resolver that combines deterministic identifier matching with LLM-assisted disambiguation for incomplete or ambiguous names.
Ligand generation uses HEALER as a molecular enumeration tool, which contains three modes of molecule enumeration. Molecule-mode is the default way of generating synthesizable analogs of a given seed molecule, performing a recursive retrosynthetic fragmentation of the query molecule using predefined reaction templates to create a retrosynthesis tree with fragment and reaction annotations. HEALER matches building blocks to fragments based on chemical similarity and recombines them via predefined reactions to produce analogs of the seed molecule. Fragment-mode allows the enumeration to directly start from fragments and generate molecules that combine the given fragments in a synthetically feasible way. Site-mode allows the reaction enumerations to take place at specified reactive sites of the seed molecule while preserving the rest. In this study, Enamine US building block stock was used by HEALER, and Molecule-mode was allowed to have at most two retrosynthesis reaction steps to prevent long enumeration runtimes.
The optimization node uses Bayesian optimization, consisting of a Gaussian process surrogate model and a qLogExpectedImprovement (qLogEI) acquisition function for single-objective optimization, and qLogNoisyExpectedHypervolumeImprovement (qLogNEHVI) for multi-objective optimization. The optimization node uses BayBE's Campaign API to recommend the next batch of molecules for evaluation. Molecules are encoded as SubstanceParameter objects using one of several fingerprint schemes: Mordred, Extended Connectivity Fingerprints (ECFP), or RDKit descriptors. The surrogate model is updated with all prior experimental results at each iteration, and the acquisition function selects the most promising unevaluated molecules. The system validates recommendations against the search space and filters out previously tested molecules before forwarding them for characterization.
Characterization backends are selected automatically based on the optimization target of interest. RDKit computes molecular descriptors such as QED, LogP, TPSA, molecular weight, H-bond donors/acceptors, rotatable bonds, ring count, etc. STOPLIGHT predicts physicochemical and structural properties, assay liabilities, and pharmacokinetic properties for small molecules. Boltz-2 accepts ligand SMILES and protein sequences (or UniProt IDs, resolved automatically) and returns affinity estimates. A 'decide characterization' node in the agent's graph inspects the target properties and selects the minimal set of tools required. Binding affinity prediction was performed using Boltz-2, a deep learning model for biomolecular structure and property prediction based on a diffusion-transformer architecture. Candidate molecules selected by the optimization module are provided to the model as SMILES strings, while the amino acid of the target protein is retrieved by the argument extraction tool. Boltz-2 jobs are dispatched asynchronously to a dedicated inference backend, either via an authenticated HTTP request to a REST endpoint in API mode, or submitted to a SLURM-managed high-performance computing cluster via a task queue in Celery mode.
Optimization loops within SABLE terminate when one of several predefined conditions is met: (i) Maximum number of iterations specified by the user has been reached; (ii) The convergence node monitors improvement in the objective function across iterations, and if no significant improvement is observed over several consecutive iterations, the system interprets this as convergence; (iii) The workflow terminates if all molecules within the enumerated search space have been evaluated; (iv) Optimization stops if the remaining unevaluated molecules are fewer than the specified batch size. SABLE is distributed with a CLI and a web-based interface designed to simplify interaction with the optimization workflow, allowing users to submit prompts describing optimization objectives, seed molecules, and monitor the progress of active optimization campaigns. User prompts may contain ambiguous or incomplete information, and if the prompt lacks certain parameters, the system uses default configurations. When critical information is missing, such as the starting molecule name or a protein target required for affinity prediction, the system fails early and informs the user of what they might need to provide.
Several caveats currently limit the agentic workflow. First, every claim SABLE makes about a candidate molecule is downstream of a predictive model, and predictive models for affinity, permeability, and toxicity remain imperfect; a library that is enriched
in Boltz-2 logIC50 is enriched against its limitations, not against experimental ground truth. Second, HEALER's synthetic accessibility is a necessary but insufficient condition for true experimental tractability; HEALER provides existing building blocks along with a reaction template, but this information alone does not constitute an experimental protocol. Furthermore, HEALER is designed to produce synthetically similar compounds; it does not provide the opportunity to assess structurally similar but synthetically dissimilar compounds. Third, the Gaussian-process surrogate scales polynomially with each evaluated compound, and while this is tractable for conservative batch sizes, very large search spaces and long campaigns will eventually require space or deep-kernel approximations.
Improvements for AI systems
Improvement 1: Domain-Delegated Agent Architecture
-
What: Implement a hard separation where the LLM only handles intent parsing, entity resolution, and result summarization—never core scientific reasoning (e.g., molecule proposal, property prediction, or optimization).
-
What the improved system can do: Eliminate hallucination risks in chemistry/physics tasks by routing all quantitative decisions to validated statistical or physics-based tools (e.g., Bayesian optimization, QSAR models, diffusion-based structure predictors). The system can now reliably execute multi-step drug-discovery campaigns without producing chemically invalid intermediates or biased property estimates.
Improvement 2: Hybrid Entity Resolution with LLM-Assisted Disambiguation
-
What: Combine deterministic regex/identifier matching with LLM fallback for ambiguous inputs (e.g., molecule names, protein identifiers, implicit objectives).
-
What the improved system can do: Accept natural-language prompts with incomplete or colloquial references (e.g.,
the kinase from the paper
orimprove brain penetration
) and resolve them to canonical SMILES/UniProt IDs with high accuracy. It can fail early with actionable feedback when critical entities are missing, reducing silent misinterpretation.
Improvement 3: Stateful, Resumable Workflow with Checkpointing
-
What: Use graph-based orchestration (e.g., LangGraph) with persistent state snapshots at every iteration, including enumerated libraries, evaluated molecules, and surrogate model history.
-
What the improved system can do: Resume interrupted optimization campaigns without recomputing expensive characterizations (e.g., Boltz-2 inference). This enables long-running, multi-day experiments on HPC clusters with fault tolerance, saving computational resources and wall-clock time.
Improvement 4: Adaptive Multi-Objective Acquisition with Pareto-Aware Exploration
-
What: Implement qLogNEHVI (or similar) for multi-objective Bayesian optimization, with early-iteration exploration to build a diverse Pareto frontier, then shift to exploitation near promising regions.
-
What the improved system can do: Simultaneously optimize conflicting objectives (e.g., potency vs. drug-likeness vs. CNS penetration) and return a set of nondominated candidate molecules rather than a single optimum. This mirrors real hit-to-lead decisions where trade-offs are necessary.
Improvement 5: Convergence-Aware Early Stopping
-
What: Monitor improvement in objective metrics across consecutive iterations and terminate when gains fall below a threshold, or when the remaining search space is smaller than batch size.
-
What the improved system can do: Reduce oracle calls by 60–90% compared to exhaustive screening, especially for large libraries (10 3–10 6 compounds). It can automatically detect diminishing returns and stop campaigns, freeing compute for other tasks.
Improvement 6: Modular Tool Backends via Configuration
-
What: Decouple characterization tools (e.g., RDKit, STOPLIGHT, Boltz-2) from the core workflow using a config-file-driven interface, allowing swap-in of new predictors without code changes.
-
What the improved system can do: Adapt to new target classes or property endpoints (e.g., replacing Boltz-2 with AlphaFold3, or adding a custom toxicity model) by editing a YAML/JSON config. This makes the system future-proof and domain-extensible beyond drug discovery (e.g., materials design, catalyst screening).
Improvement 7: Synthetic-Tractability-Aware Enumeration
-
What: Integrate reaction-template-based enumeration (e.g., HEALER) with building-block availability and retrosynthetic depth limits (e.g., max 2 steps).
-
What the improved system can do: Propose analogs that are not only chemically valid but also synthesizable from commercial building blocks, reducing the gap between computational hits and experimentally testable compounds. It can also flag when a desired analog requires unavailable reaction templates, as seen in the METTL3 case.
Improvement 8: Uncertainty-Aware Reporting
-
What: Propagate predictive model uncertainties (e.g., from Boltz-2 or QSAR) into the final candidate ranking and Pareto front, and explicitly state that enrichment is relative to model limitations, not experimental ground truth.
-
What the improved system can do: Provide decision-makers with confidence intervals on predicted affinities/properties, preventing overcommitment to false positives. It can also prioritize candidates with high predicted improvement and low uncertainty, improving experimental hit rates.
Improvement 9: Scalable Surrogate Modeling for Large Search Spaces
-
What: Replace standard Gaussian processes with sparse or deep-kernel approximations when the enumerated library exceeds 10 5 compounds or when iterations grow large.
-
What the improved system can do: Maintain tractable optimization runtimes for million-compound libraries and long campaigns (e.g., >100 iterations), without sacrificing acquisition quality. This enables broader exploration of chemical space.
Improvement 10: Automated Objective Detection and Default Configuration
-
What: Use the LLM to infer the number and type of objectives (e.g.,
minimize binding affinity and maximize QED
) from the prompt, then auto-select the appropriate acquisition function and characterization tools. -
What the improved system can do: Allow non-expert users to run complex multi-objective optimizations by simply describing their goals in natural language. The system fills in missing parameters (e.g., batch size, iteration limit) with sensible defaults, lowering the barrier to entry for medicinal chemists.
Abstract
Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration), an open-source framework that employs natural-language orchestration to guide chemical structure optimization. SABLE uses an LLM to interpret user-defined goals and route tasks, while specialized tools perform reaction-templated analog enumeration, physicochemical and ADMET property prediction, structure-based affinity scoring, and Bayesian optimization. The resulting workflow is a computational twin of the analytical and prioritization stages of the design-make-test-analyze cycle, providing provenance of each numerical output. Across single, and multi-objective optimization studies, SABLE enriches candidate sets for user-defined computational objectives while evaluating only a subset of the enumerated search space. Its modular architecture allows tools and characterization backends to be replaced by editing a simple config file, without modifying operational logic. SABLE provides an extensible decision-support framework for prioritizing synthetically constrained analogs in early-stage drug discovery.
Sources
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection