Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Beyond Scaffold Splits".
Jane: Molecular property models are routinely evaluated by holding out Bemis–Murcko scaffolds, yet a scaffold label captures only one notion of chemical novelty.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap, this new paper is showing that relying on standard scaffold splits in ADMET model testing misses a massive gap by not accounting for molecules that are both structurally remote and physically sparse from what the AI has seen.
Jane: Exactly, and what’s really significant is how much error this frontier exposes; they found a median relative error inflation of eighty-seven percent across most tasks, with some specific predictions even completely flipping their ranking on things like blood-brain barrier permeability.
Lu: The real excitement here is the methodology they developed for this structural-frontier split; it’s not just a simple scaffold label anymore, but something that actively identifies the sparsest and most remote groups based on descriptor radius and local sparsity.
Meng: From my side, I'm focusing on the practical takeaway: if models are failing so badly when they see these truly novel structures, it tells us that simply scaling up model size isn't the answer; we need to fundamentally rethink how we curate our training data to cover those sparsely populated regions better.
Lalam: I think this research has huge implications for how we build and trust these systems; it suggests that our current evaluation pipelines are incomplete, and this paper gives us a much more rigorous way to stress-test the models against real-world novelty risk.
Tom: It’s definitely a call for better rigor in our testing protocols, Jane, because they showed that even advanced robust objectives didn't reliably close the gap when facing these frontier groups.
Jane: That’s a tough pill to swallow, Tom; it means we can’t just rely on tweaking our risk penalties to fix generalization issues in this specific structural sense.
Lu: But the authors make a very strong point about the distinction between data support and objective design, suggesting that we need to improve the chemical environment definition first before we try to build perfect risk measures for it.
Meng: So, what’s next on our agenda? We saw they compared this frontier gap against established splits like Lo-Hi and DataSAIL, which was a really interesting comparison point.
Lalam: That comparison is crucial because it showed that different splits expose different weaknesses in the models; you can't just swap them out blindly for a better one.
Tom: Precisely, and the final conclusion they landed on is that we should report these structural frontiers alongside existing splits instead of trying to replace them entirely.
Jane: It sounds like the big message here is that we need a more holistic view of chemical novelty risk when evaluating AI models for property predictions.
The paper's summary: Tom: So, to wrap up the paper's main points, they’re suggesting we stop treating these scaffold splits as the final word and instead incorporate this structural frontier testing alongside them for a more complete picture of model performance.
Jane: That’s right, and the authors are proposing that instead of just picking one split over another, we need to look at how models behave across all these different ways we partition chemical space.
Lu: They are pushing for this structural-frontier approach because it captures a much more complete picture of where the model is genuinely failing, especially when dealing with those remote and sparse scaffold groups.
Meng: From an engineering standpoint, this means our future testing frameworks should probably be designed to handle multiple, different types of structural stress tests simultaneously rather than just one standard check.
Lalam: I think the implication for our culture is that we need to foster a mindset where we don't settle for the first evaluation method we find; instead, we embrace complexity in our testing to ensure true reliability.
Tom: They also pointed out that while they tried various robust objectives, none of them managed to reliably close that gap on their own, which really highlights the data support problem before it becomes an objective design problem.
Jane: So, they're saying we shouldn't waste too much time tuning our risk measures when the fundamental issue is how we are selecting and labeling the training data to better cover those difficult chemical environments.
Lu: That makes sense; if you feed a model data that’s already sparse in certain regions, no amount of penalty tuning will make it perform well on novel chemistry.
Meng: I see what they mean about the sequencing of priorities; we need to prioritize improving the quality and coverage of our datasets for these remote scaffolds before we worry about optimizing every single robustness objective.
Lalam: This shifts our focus toward data curation, which is a really important area because it means the quality of our foundational models depends heavily on how well we cover these sparsely populated regions upfront.
The paper's improvements: Tom: So, to wrap up this deep dive into "Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models," the main message is that we need a much broader view when testing AI for these kinds of property predictions.
Jane: Exactly, it shows us that standard scaffold splits aren't enough; they only cover a tiny slice of the chemical landscape where models might actually break down.
Lu: The potential here is huge because this opens up entirely new ways to define what constitutes a challenging or novel structure for our AI to handle, which could lead to much more robust predictive tools in drug discovery.
Meng: From an engineering standpoint, this means our future testing pipelines should probably be designed to handle multiple, different types of structural stress tests simultaneously rather than just one standard check.
Lalam: I think the real cultural impact is seeing the commitment to this level of rigor; it pushes us all toward a more thoughtful approach to how we trust and deploy these powerful AI systems in critical areas.
Tom: They concluded that we should report these structural frontiers alongside existing splits instead of trying to replace them entirely, which is a really practical suggestion for our daily work.
Jane: That's smart because it gives us the most complete picture possible about where the model is succeeding and where it’s falling short, rather than just giving us one number.
Lu: I think this research paves the way for creating more adaptive testing strategies that can dynamically adjust their stress tests based on how a model performs across different chemical views.
Meng: That sounds complex to build, but if we can make those dynamic tests manageable, it could significantly reduce the risk of unexpected failures when we deploy these models in real applications.
Lalam: It really emphasizes that the goal isn't just to build better predictors, but to build a whole system of evaluation that is truly comprehensive and honest about its limits.
Conclusion: Tom: So, we’ve just finished diving deep into "Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models," and what we found is that standard scaffold splits are leaving massive gaps in how we test these AI models for drug property predictions.
Jane: It really shows us that relying on just a few ways to group molecules isn't enough if we want to know if our models are truly reliable across the whole chemical universe.
Lu: The potential here is huge because this opens up entirely new ways to define what constitutes a challenging or novel structure for these AI to handle, which could lead to much more robust predictive tools in drug discovery.
Meng: From an engineering standpoint, this means our future testing frameworks should probably be designed to handle multiple, different types of structural stress tests simultaneously rather than just one standard check.
Lalam: I think the real cultural impact is seeing the commitment to this level of rigor; it pushes us all toward a more thoughtful approach to how we trust and deploy these powerful AI systems in critical areas.
Tom: They concluded that we should report these structural frontiers alongside existing splits instead of trying to replace them entirely, which is a really practical suggestion for our daily work.
Jane: That's smart because it gives us the most complete picture possible about where the model is succeeding and where it’s falling short, rather than just giving us one number.
Lu: I think this research paves the way for creating more adaptive testing strategies that can dynamically adjust their stress tests based on how a model performs across different chemical views.
Meng: That sounds complex to build, but if we can make those dynamic tests manageable, it could significantly reduce the risk of unexpected failures when we deploy these models in real applications.
Lalam: It really emphasizes that the goal isn't just to build better predictors, but to build a whole system of evaluation that is truly comprehensive and honest about its limits.
Tom: So, "Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models" shows us that we need a much broader view when testing AI for these kinds of property predictions.
Jane: Exactly, it shows us that standard scaffold splits aren't enough; they only cover a tiny slice of the chemical landscape where models might actually break down.
Lu: The potential here is huge because this opens up entirely new ways to define what constitutes a challenging or novel structure for these AI to handle, which could lead to much more robust predictive tools in drug discovery.
Meng: From an engineering standpoint, this means our future testing frameworks should probably be designed to handle multiple, different types of structural stress tests simultaneously rather than just one standard check.
Lalam: I think the real cultural impact is seeing the commitment to this level of rigor; it pushes us all toward a more thoughtful approach to how we trust and deploy these powerful AI systems in critical areas.
Tom: They concluded that we should report these structural frontiers alongside existing splits instead of trying to replace them entirely, which is a really practical suggestion for our daily work.
Jane: That's smart because it gives us the most complete picture possible about where the model is succeeding and where it’s falling short, rather than just giving us one number.
Lu: I think this research paves the way for creating more adaptive testing strategies that can dynamically adjust their stress tests based on how a model performs across different chemical views.
Meng: That sounds complex to build, but if we can make those dynamic tests manageable, it could significantly reduce the risk of unexpected failures when we deploy these models in real applications.
Lalam: It really emphasizes that the goal isn't just to build better predictors, but to build a whole system of evaluation that is truly comprehensive and honest about its limits.
Tom: Our deep dive into "Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models" shows us that standard scaffold splits miss a huge problem where models fail spectacularly on molecules that are both remote and unsupported.
Jane: It really highlights that the way we partition chemical space matters deeply for model performance, especially when dealing with complex properties like ADMET predictions.
Lu: The paper’s suggestion is to adopt a structural-frontier split, focusing on the most sparse and remote groups, because that seems to be where the real challenge lies for these property models.
Meng: Practically speaking, this means we need to rethink how we define our applicability domains so they account for this specific type of structural novelty risk.
Lalam: This research gives us a clearer picture of the necessary rigor needed in our evaluation pipelines to ensure our AI systems are truly reliable across the entire chemical universe.
Jiacheng Zheng, Chang Guo, Zixuan Wang, Xinyu Liu, Hao Chen
Marine College, Shandong University · Department of Mathematics, Faculty of Mathematical & Physical Sciences, University College London Department of Mathematics, Faculty of Mathematical & Physical Sciences, University College London School of Life Science and Technology, Harbin Institute of Technology School of International Trade and Economics, University of International Business and Economics
cs.LG, q-bio.QM
Submitted: 2026-07-12
Updated: 2026-10-06
Importance score: 87/100
The gist: Molecular property models are routinely evaluated by holding out Bemis–Murcko scaffolds, yet a scaffold label captures only one notion of chemical novelty.
Key concepts
- Bemis–Murcko Scaffold Split
- This common method groups molecules based on a shared ring and linker structure. It is used to test models but can fail because it forces a continuous chemical space into a single, rigid binary partition, potentially missing diverse structural challenges.
- Structural-Frontier Split
- This new evaluation reserves the 'sparsest and most physicochemically remote scaffold groups.' This split tests models on chemistry that is both chemically distant and locally unsupported, exposing failures in generalization beyond familiar structural neighborhoods.
- Error Inflation
- This measures how much worse a model performs when tested on the frontier compared to a standard control. The study found a median error increase of 87% for equally weighted endpoints, indicating that models struggle significantly when presented with novel, distant chemical structures.
- Message-Passing Graph Network Control
- This is a higher-capacity method used to evaluate model performance. It involves using a graph network to process information between molecules, which helps determine if the observed error inflation on the frontier is a genuine structural challenge or merely an artifact of limited model capacity.
Terminology
Summary
Molecular property models are routinely evaluated by holding out Bemis–Murcko scaffolds, yet a scaffold label captures only one notion of chemical novelty. The frontier inflates primary error by a taskwise median of 87.0% (skew-sensitive mean 130.3%); the gap survives a message-passing graph network and, on average, exceeds the published Lo-Hi and DataSAIL splits.
The gist
A structural-frontier holdout increases equally weighted primary error by a median of 87.0% (skew-sensitive mean 130.3%) relative to a ratio- and group-matched scaffold control, and this gap survives a higher-capacity graph-network encoder while exceeding published hard splits.
The Problem Addressed
Machine learning models for ADMET are often evaluated using methods like the Bemis–Murcko scaffold split, which mitigates leakage by grouping molecules around a common ring-and-linker framework. However, this approach converts a continuous and representation-dependent support question into a single binary partition. Recent evaluations show that molecular ranking can change substantially with the dataset, representation, and split; simulated prospective splits, low-similarity benchmarks, distance-aware evaluation methods, and leakage-controlled splitting each expose failure modes that an ordinary scaffold holdout may miss. The study questions whether a model generalizes to a structural frontier that is simultaneously remote and locally unsupported, and whether this answer survives different chemically defensible representations.
The Methodology
The authors introduce a label-free structural-frontier split based on scaffold-level robust descriptor radius and local sparsity,
which reserves the sparsest and most physicochemically remote scaffold groups.
This frontier is benchmarked against a ratio- and group-matched scaffold control, retaining the conventional 80/10/10 split as context. The study also formulates MV-FREX as a falsifiable probe that penalizes a count-adjusted tail and outward risk slope across ECFP, MACCS, atom-pair and physicochemical views.
The comparison involves 468 fully recorded primary runs with five paired seeds, fixed robust-penalty controls, ablations, calibration metrics, linear controls (a four-run message-passing graph network control), and an empirical comparison against the Lo-Hi and DataSAIL splitters.
Key Findings on Error Inflation
The frontier exposes a large endpoint-wide gap across all tasks. Across equally weighted endpoints, the taskwise median relative error inflation over the ratio- and group-matched scaffold control is 87.0%, with a skew-sensitive mean of 130.3%. This effect is dominated by one endpoint: for blood–brain-barrier permeability (BBB), AUROC falls from 0.879 ± 0.020 to 0.409 ± 0.005, representing a genuine score-ranking inversion, not a prevalence artifact. The gap survives a higher-capacity encoder, as the mean normalized error inflation is 82.8% across four tasks when using a message-passing graph network control.
Robustness and Limitations of Controls
The study investigates whether specific robust objectives can reliably close the gap. A count-adjusted multi-view tail-risk penalty and three fixed distributionally robust objectives do not reliably close the gap.
The V-REx-style, smooth worst-group, and sample-tail CVaR-style penalties were also found to be mixed,
with every interval including zero. The authors conclude that robustness at the frontier is a data-support problem before it is an objective design problem, as environment penalties protect the environments they can see rather than unseen chemistry. Furthermore, pairwise Spearman correlations between view-specific novelty scores range from −0.093 to 0.837 (mean 0.321), supporting a multi-view evaluation but not the claim that optimizing the worst source-view risk controls unseen chemistry.
Comparison with Existing Splits
The matched frontier gap is robust to model capacity and exceeds two published hard splits, Lo-Hi and DataSAIL. The frontier is on average the most severe (mean error inflation over matched control is 118.9% for the frontier versus 96.0% for DataSAIL and 54.1% for Lo-Hi). Crucially, the splitters are not interchangeable: DataSAIL makes HIA harder and PPBR easier, while Lo-Hi could not form a low-similarity PPBR partition at all. This suggests that each stress test exposes a different axis of chemical space, leading the authors to recommend reporting the frontier beside
existing splits rather than as a replacement. Finally, CMNPD audit results show that both OOD status and legacy-teacher agreement vary with representation and endpoint; predicted database annotations are not substitutes for experimental validation.
Improvements for AI systems
Here are specific improvements that can be made to AI systems by leveraging the findings of this research:
-
Improve prospective compound prioritization by moving beyond simple scaffold holdouts to a
structural-frontier split
that explicitly reserves groups that are both physically sparse and chemically remote from the training data, rather than just those with similar ring-and-linker frameworks. -
Enhance model reliability assessment by incorporating a multi-view ordering check (ECFP, MACCS, atom pair, physicochemical) as a mandatory constraint during split construction and model evaluation. This prevents models from appearing robust on one descriptor family while failing on another.
-
Develop more rigorous robustness testing for training penalties by moving away from fixed robust penalties (like standard CVaR or V-REx-style methods) toward objectives that are demonstrably tied to the specific chemical environments being tested, as current objectives do not reliably close the frontier gap across different model architectures.
-
Refine Applicability Domain (AD) declarations by treating them as a deterministic stress test based on ratio- and group-matched allocation, rather than relying on static confidence scores or heuristic domain cutoffs. This allows for the creation of more scientifically defensible
applicability domains
that reflect actual structural novelty risk. -
Improve model interpretation by specifically analyzing ranking inversions in novel chemical space (like the BBB effect). Systems should be capable of identifying when a model's score reverses its relationship with the true label as it approaches sparse regions, flagging this specific failure mode rather than just reporting an overall error inflation metric.
-
Enhance data provenance auditing by treating computational annotations (like those from legacy models) with skepticism and only using them for teacher agreement audits, while reserving
accuracy
claims strictly for experimentally measured endpoints. This prevents the dilution of biological ground truth by noisy legacy predictions in OOD evaluations.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks