Periodic Topological Deep Learning for Polymer Design and Discovery

arXiv:2605.26833 · cs.LG, cs.AI · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Periodic Topological Deep Learning for Polymer Design and Discovery".

Jane: The paper was written by Yasharth Yadav, Tze Kwang Gerald Er, Atsushi Goto and Kelin Xia from Nanyang Technological University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we're digging into a paper that just landed on arXiv, and it's called "Periodic Topological Deep Learning for Polymer Design and Discovery." Jane, I have to say, the title alone made me sit up a little straighter.

Jane: Same here, Tom. And honestly, the first thing that struck me is the word "periodic." We've seen a lot of machine learning work on polymers before, but most of it treats a polymer like it's just a small molecule. You take one repeating unit, you draw it as a graph, and you call it a day. This paper says that's actually a pretty big blind spot.

Tom: Right, because a polymer isn't a single unit. It's this long, repeating chain that can stretch for millions of atoms. If you only look at one unit, you miss the connections between units, and you definitely miss the way the chain folds and interacts with itself.

Jane: Exactly. And that's where the "topological" part comes in. The authors, Yasharth Yadav and the team at NTU, they're using something called a Vietoris–Rips complex. I know that sounds like a mouthful, but the idea is pretty elegant. Instead of just drawing lines between atoms that are bonded, you draw connections between any atoms that are close together in space.

Tom: So it's not just about covalent bonds anymore. It's about capturing those many-body interactions, where three or four atoms are all within a certain distance of each other at the same time. That's something a regular graph just can't do.

Jane: Right. A graph can tell you that atom A is connected to atom B, but it can't easily tell you that atoms A, B, and C are all mutually interacting in a triangle. That triangle is a higher-order structure, and it carries real physical meaning, especially when you're thinking about things like hydrogen bonding or steric hindrance.

Tom: And the "periodic" part means they're doing this across the whole chain, not just within one monomer. They actually build a distance matrix that accounts for the fact that the chain repeats, so atoms at the boundary of your chosen unit are correctly seen as neighbors to atoms in the next unit.

Jane: That's the part that got me excited. It's a representation that finally respects what a polymer actually is. It's not a small molecule, it's a periodic structure. And they built a whole deep learning framework around that idea, which they call Periodic-TDL.

Tom: And the results, we'll get into those in a bit, but let me just tease this: they beat every state-of-the-art model they compared against on eight out of nine property prediction tasks. That's a strong opening statement.

Jane: It is. And the implications go beyond just benchmarks. If you can predict properties like glass transition temperature more accurately, you can start designing new polymers for batteries, for medical devices, for packaging, without having to synthesize hundreds of candidates in the lab first.

Tom: So the title really is doing a lot of work here. "Periodic" for the chain structure, "Topological" for the higher-order interactions, and "Deep Learning" for the encoder that learns from it all. We'll unpack each of those pieces as we go.

Jane: And we should also mention that they didn't just stop at prediction. They went and synthesized three brand new polymers to test their model's predictions. That's the kind of validation that really makes a paper stand out.

Tom: Absolutely. We'll get to that in the next segment, where we break down the actual methodology and how they built this thing. Stick around.

Summary: Tom: So we're back, still talking about "Periodic Topological Deep Learning for Polymer Design and Discovery." Jane, we teased the methodology, so let's actually dig into how this thing works.

Jane: Okay, so the core idea is that they take the polymer and they run a filtration. That's a fancy way of saying they look at the structure at different distance cutoffs. At a very small cutoff, say two angstroms, you only see covalent bonds. That's your standard molecular graph.

Tom: But then they increase the cutoff to three angstroms, and then to four. At those larger scales, you start picking up non-covalent interactions. Atoms that aren't bonded but are close in space, maybe through hydrogen bonding or just van der Waals contacts. And at each of those scales, you get a simplicial complex, which is a graph plus triangles plus even higher-dimensional shapes.

Jane: Right. And the clever part is how they learn from all of this. They built something called a Hierarchical Simplicial Message Passing encoder, or HSMP. The idea is that information flows from the coarsest scale, the four angstrom cutoff, down to the finest scale, the two angstrom covalent graph.

Tom: So the long-range interactions inform the short-range ones. The model learns to use the big picture to refine the local picture. That's a really nice design choice, because it means the final atom representations carry both local chemical context and global structural context.

Jane: Exactly. And they also did something I haven't seen before in this space. They used Forman–Ricci curvature as a feature for the simplices. That's a geometric measure that tells you something about the local shape of the structure. Positive curvature might indicate a dense region, negative curvature might indicate a bottleneck.

Tom: And that's not just a gimmick. It gives the model a principled way to initialize features for those higher-order simplices, the triangles and beyond. Without that, you'd have no idea what to put in there.

Jane: Right. So they pretrain this encoder on a dataset called PI1M, which is a million unlabelled polymers. They use self-supervised tasks like predicting atom contexts and functional groups. Then they fine-tune on nine specific property prediction tasks.

Tom: And the results, Jane, they're really something. On the glass transition temperature task, which is experimentally measured, they got a root mean square error of about thirty-four point seven degrees Celsius. The standard deviation of that dataset is over one hundred ten degrees. So they're cutting the error by about sixty-nine percent relative to just guessing the mean.

Jane: And they beat the best baseline on eight of nine tasks. The only one they didn't win outright, they still tied for the best. That's a clean sweep, essentially.

Tom: But here's the thing that really impressed me. They didn't just stop at benchmarks. They used their model to study a specific chemical question. They generated a dataset of over forty-eight thousand polymers by systematically substituting groups on acrylate and acrylamide backbones.

Jane: And they looked at two specific modifications. Replacing an ester group with an amide group, and adding a methyl group to the backbone. Both are known to affect glass transition temperature, but this paper quantified the effects across thousands of matched pairs.

Tom: The ester-to-amide switch gave a mean increase of about fifty-five degrees Celsius. The backbone methylation gave about fourteen degrees. And both were highly consistent across the matched pairs.

Jane: And then they went and checked against real experiments. They synthesized three polymers that had never been made before, and they pulled four more pairs from the literature. The directional predictions matched in all six cases.

Tom: That's the kind of validation that makes you trust the model. It's not just fitting noise, it's capturing real physics. And that's a huge deal for polymer design.

Jane: It really is. So we've got the representation, we've got the learning, and we've got the validation. Next we should talk about what this means for the field and where it might go from here.

Improvements: Tom: We're back with "Periodic Topological Deep Learning for Polymer Design and Discovery." Jane, we've covered the basics and the results. Now let's talk about what this paper actually improves upon and why it matters.

Jane: So the biggest improvement, I think, is the representation itself. Most existing models, even the good ones, treat a polymer as a single repeating unit. They build a graph of that one unit and call it done. This paper says that's fundamentally incomplete.

Tom: Right, because you're missing the periodicity. The chain extends, and the interactions between adjacent units matter. The authors built a periodic distance matrix that accounts for the translational symmetry of the polymer. That alone is a significant step forward.

Jane: And then there's the move from graphs to simplicial complexes. Graphs can only encode pairwise interactions. A simplicial complex can encode three-body, four-body, even higher-order interactions. For polymers, where things like hydrogen bonding networks are critical, that extra expressive power really matters.

Tom: And the way they learn from it, that hierarchical message passing, is also an improvement. Instead of treating each filtration scale independently and just concatenating the features at the end, they propagate information from coarse to fine. The covalent bond features at the finest scale are enriched by the long-range interactions from the coarser scales.

Jane: That's a design choice that has a real physical motivation. The local environment of an atom is influenced by the global structure of the chain. By forcing the information to flow that way, the model is learning representations that are more chemically meaningful.

Tom: And they also introduced curvature-based features, which is new for topological deep learning in this context. Forman–Ricci curvature gives you a geometric signal that's scale-dependent, which fits perfectly with the filtration approach.

Jane: Let me also mention the practical side. Their encoder, at the finest scale, reduces to a standard molecular graph. That means it can plug directly into existing pretraining pipelines. You don't have to reinvent the wheel for self-supervised learning. That's a huge practical advantage.

Tom: And the results speak for themselves. On the glass transition temperature task, they got an R-squared of about zero point nine zero. That's excellent for a property that's notoriously hard to predict. And they did it while also beating baselines on electronic properties like bandgap and ionization energy.

Jane: The ablation studies are also worth mentioning. They showed that removing the periodic representation hurts performance. Removing the hierarchical message passing hurts performance. Removing the multi-head attention hurts performance. Every component is pulling its weight.

Tom: So what does this mean for the field? I think it means that the era of treating polymers as small molecules is over. If you want to do polymer informatics properly, you need to respect the periodic, many-body nature of the material.

Jane: And I think it also opens the door for more ambitious design tasks. If you can predict properties this accurately, you can start doing inverse design. You can ask the model to generate polymers with a specific glass transition temperature, or a specific dielectric constant, and it can guide you toward promising candidates.

Tom: They even showed a connection between predicted glass transition temperature and synthetic accessibility. Higher Tg polymers tend to be bulkier and more polar, which makes them harder to synthesize. That kind of insight is really useful for experimentalists.

Jane: So the improvements here are not just incremental. They're foundational. The representation is better, the learning is better, and the validation is more rigorous. This is a paper that could really shift how the community approaches polymer informatics.

Tom: And we should note, they've made all the code and data publicly available. So anyone can build on this. That's how science moves forward.

Jane: Absolutely. So let's wrap up and think about the big picture. What does this mean for the world?

Conclusion: Tom: Alright, we're at the end of our time with "Periodic Topological Deep Learning for Polymer Design and Discovery." Jane, let's pull it all together.

Jane: So the big picture here is that this paper gives the polymer informatics community a better tool. It's not just a slightly better predictor, it's a fundamentally more faithful representation of what a polymer actually is. Periodic, many-body, multiscale.

Tom: And that faithfulness translates directly into better predictions. They beat every baseline on nearly every task, and they validated their predictions with real experiments, including three polymers that had never been synthesized before.

Jane: The experimental validation is what really sets this apart. They predicted that ester-to-amide substitution would raise the glass transition temperature by about fifty-five degrees, and the experiments confirmed the direction. They predicted that backbone methylation would raise it by about fourteen degrees, and again, the experiments agreed.

Tom: That's not just a model that fits training data. That's a model that has learned something real about chemistry. And that's the kind of tool that can accelerate materials discovery in a meaningful way.

Jane: Think about the applications. Better polymer property prediction means faster development of new materials for batteries, for medical implants, for sustainable packaging, for electronics. Every time you can predict a property without synthesizing a hundred candidates, you save time and resources.

Tom: And the fact that they made everything open source means the whole community can build on this. Future work could extend it to copolymers, to crosslinked networks, to more complex architectures. The foundation is solid.

Jane: There are limitations, of course. It's currently restricted to linear homopolymers. It doesn't handle crosslinking. But those are directions for future work, not flaws in what's been done.

Tom: And honestly, for a first framework that combines periodicity and topology in polymer deep learning, this is an impressive debut. The authors should be proud of what they've built.

Jane: Agreed. So we'll say goodbye to this paper, and we'll be back soon with the next one. Thanks for listening, everybody.

Tom: Take care, and keep learning.

Yasharth Yadav, Tze Kwang Gerald Er, Atsushi Goto, Kelin Xia

Nanyang Technological University · Nanyang Technological University

cs.LG, cs.AI

Submitted: 2026-08-16

Updated: 2026-08-18

Comments: 19 pages, 3 figures, 3 tables

Code: https://github.com/yasharthy/Periodic-TDL

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

Key concepts

Periodic Structure
This concept acknowledges that a polymer is not a single molecule but a long, repeating chain. The model accounts for this translational symmetry by building a distance matrix that correctly identifies atoms at the boundary of chosen units as neighbors to atoms in the next unit.

Terminology

Summary

Summary

This paper introduces Periodic-TDL, a topological deep learning framework for polymer property prediction and discovery. The authors identify a fundamental representational bottleneck in existing polymer machine learning methods: most approaches represent polymers as molecular graphs of a single repeating unit, thereby missing both the periodicity of polymer chains and many-body interactions beyond pairwise bonds. To address this, they propose a framework built on periodic Vietoris-Rips complexes that capture many-body interactions across multiple spatial scales, followed by a Hierarchical Simplicial Message Passing (HSMP) encoder.

The core methodological contributions are as follows. First, the authors construct a periodic distance matrix by incorporating translated copies of the repeating unit via iterative rearrangement of backbone fragments while strictly preserving chemical connectivity, defining each interatomic distance as the minimum separation across all translated copies. This restores the correct spatial proximity between atoms belonging to adjacent repeating units and ensures invariance to the arbitrary choice of repeating unit. From this periodic distance matrix, a Vietoris-Rips filtration is applied at three cutoff distances: ϵ1 = 2.0 Å, ϵ2 = 3.0 Å, and ϵ3 = 4.0 Å. The complex at ϵ1 operates at the scale of covalent bonds, while ϵ2 and ϵ3 progressively incorporate non-covalent interactions. The resulting simplicial complexes encode both covalent and non-covalent interactions through higher-dimensional simplices.

Second, the authors develop the HSMP encoder, which operates over the nested sequence of simplicial complexes and propagates information from coarser to finer scales in a chemically motivated direction. The encoder combines two update functions: (i) multi-head simplicial message passing, which extends the Message Passing Simplicial Network framework to a multi-head setting, allowing multiple parallel simplex updates operating on different representation subspaces; and (ii) cross-scale refinement, which enriches finer-scale simplex representations using long-range interaction information from coarser scales through a residual gated mechanism. Message passing proceeds hierarchically from the coarsest to the finest scale, and at the finest filtration level (ϵ1 = 2.0 Å), the encoder reduces to a standard molecular graph, enabling direct integration into established self-supervised pretraining pipelines.

Third, the authors introduce curvature-based featurization using Forman's discretization of Ricci curvature, which generalizes naturally to simplices of any dimension. For each filtration level, curvature values are computed over a finer set of Vietoris-Rips complexes around each base cutoff (with increments of 0.25 Å), yielding five curvature values per simplex. Curvature values are normalized using a temperature-scaled sigmoid transformation followed by centering to the interval (−1, 1). Initial simplex features combine chemical descriptors (atom and bond features computed using RDKit) with curvature-based geometric features.

The authors pretrained the HSMP encoder on the PI1M dataset of one million polymers using three self-supervised tasks adapted from the GROVER framework: atom context prediction, bond context prediction, and functional group prediction. The pretraining objective was a weighted sum of the three losses with weights w atom = 2.0, w bond = 1.0, and w fg = 5.0. The model was optimized using AdamW with learning rates of 2 × 10−4 for the encoder and 10−3 for task-specific prediction heads, trained for 10 epochs with a batch size of 64. The selected checkpoint achieved a training loss of 0.0763 and a validation loss of 0.0593.

The framework was finetuned on nine downstream polymer property prediction tasks spanning electronic (Egc, Eib, Egb, Eea, Ei), optical (EPS, Nc), physical (Xc), and thermal (Tg) properties. Here, Tg is experimentally measured and the remainder are DFT-derived. Finetuning followed a two-stage protocol: first, the encoder was frozen and only the regression head was optimized for 10 epochs; second, both encoder and head parameters were jointly optimized for 60 epochs with a cosine learning rate schedule with periodic restarts. Performance was assessed via five-fold cross-validation and reported as RMSE and R2 on test folds.

Periodic-TDL achieved the lowest RMSE on eight of nine tasks and the highest R2 on all nine, with consistent gains across all property types. For eight of nine targets, mean RMSE of Periodic-TDL was 50–75% lower than the standard deviation of the dataset. Specifically, for the Tg dataset, RMSE was approximately 35 °C against a standard deviation of 110 °C (69% reduction). Normalized RMSE ranged from 4–7% for most targets, 8% for EPS and 20% for Xc. Ablation studies examining (i) the role of periodic representation and hierarchical message passing, and (ii) the effect of multi-head message passing, pretraining, and model capacity, suggest that performance degrades when any component is removed.

Beyond predictive accuracy, the authors assessed the chemical credibility of Periodic-TDL by examining changes in glass transition temperature (Tg) upon ester-to-amide substitution and backbone α-methylation in acrylate- and acrylamide-based polymers. They generated a computational dataset of 48,208 monomer structures via systematic substitution on the phenyl ring of four monomers: 2-(acryloyloxy)ethyl benzoate (Ar–Et–A), 2-(methacryloyloxy)ethyl benzoate (Ar–Et–MA), 2-acrylamidoethyl benzoate (Ar–Et–AM), and 2-methacrylamidoethyl benzoate (Ar–Et–MAM). Systematic substitution was performed at ortho, meta, and para positions with four substitution patterns (no substituent, mono-, di-, and tri-substitution) and thirteen functional groups including alkyl and halogens, yielding 12,052 unique monomers per family.

The predicted Tg distributions showed statistically significant differences across all four families (p < 0.001). Pairwise analyses over 12,052 matched polymer pairs revealed that ester-to-amide substitution produced the largest and most consistent Tg elevation: the predicted shift was positive across all matched pairs for both the acrylate–acrylamide comparison (mean = 57.0 °C, 99% CI: [56.6, 57.3]) and the methacrylate–methacrylamide comparison (mean = 54.6 °C, 99% CI: [54.2, 55.0]). Backbone α-methyl substitution produced a smaller but statistically significant elevation in both the acrylate–methacrylate (mean = 15.4 °C, 99% CI: [15.3, 15.6]) and the acrylamide–methacrylamide comparisons (mean = 13.0 °C, 99% CI: [12.8, 13.2]), with a minority of matched pairs exhibiting negative ΔTg pred in both cases. All four comparisons were statistically significant (p < 0.001, two-sided t-test, corrected for multiple comparisons).

To validate these trends experimentally, the authors synthesized three polymers absent from the fine-tuning dataset and previously unreported in the literature: poly(2-(acryloyloxy)ethyl benzoate) (PBz-Et-A), poly(2-(methacryloyloxy)ethyl benzoate) (PBz-Et-MA), and poly(2-acrylamidoethyl benzoate) (PBz-Et-AM). Two of the three predictions agree with experiment to within 4 °C (PBz-Et-A: experimental 7.3 °C vs. predicted 7.4 ± 4.5 °C; PBz-Et-MA: experimental 32.7 °C vs. predicted 29.1 ± 4.4 °C). The third polymer (PBz-Et-AM) showed a larger discrepancy (experimental 38.6 °C vs. predicted 60.3 ± 5.7 °C). The three synthesized polymers yielded two matched pairs, and these were supplemented with four additional polymer pairs drawn from the experimental literature, all absent from the fine-tuning dataset. Predicted directional trends were in agreement with experiment in all six cases, spanning both ester-to-amide substitution and α-methylation effects across structurally diverse pendant group chemistries.

The authors discuss the mechanistic origins of the observed trends. The higher Tg of (meth)acrylamides relative to (meth)acrylates is attributed to two mechanisms: (i) the partial double-bond character of the C–N bond increases pendant-group stiffness in (meth)acrylamides, whereas the C–O bond in (meth)acrylates allows free rotation; and (ii) the acrylamide group can act as both a hydrogen-bond donor and acceptor, enabling physical cross-linking between amide and carbonyl groups. The higher Tg of methacrylates (methacrylamides) relative to acrylates (acrylamides) is attributed to steric hindrance from the α-methyl group, which restricts chain mobility. The authors speculate that the agreement with known trends likely arises from two features of Periodic-TDL: (i) the periodic Vietoris–Rips filtration at larger cutoffs (ϵ2 = 3.0 Å and ϵ3 = 4.0 Å) incorporates interactions at length scales relevant to non-covalent contacts such as hydrogen bonding; and (ii) the additional α-methyl group increases local steric density, leading to more higher-dimensional simplices at intermediate cutoffs, which is likely reflected in curvature-based simplex features where more positive Forman–Ricci curvature indicates locally dense regions.

The authors also note several limitations of the current work: (i) Periodic-TDL is currently restricted to linear homopolymers and does not yet handle copolymers; (ii) the periodic Vietoris–Rips complex is constructed from cyclic permutations of a single repeating unit and does not explicitly encode interactions between atoms in distant repeat units; (iii) crosslinking is not represented in the current framework; and (iv) the computational analysis suggests that higher-Tg polymers tend to be bulkier and more polar, a trend that could be tested systematically through targeted synthesis.

The paper concludes that Periodic-TDL captures the underlying physical effects of specific functional group modifications, rather than merely optimizing predictive performance on benchmark datasets. All code and data are publicly available at https://github.com/yasharthy/Periodic-TDL.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:


Improvement: Replace standard graph-based polymer encoders (which treat polymers as finite molecular graphs) with a Periodic Vietoris–Rips Complex encoder that:

  • Constructs a periodic distance matrix via cyclic permutation of backbone fragments (preserving chemical connectivity).

  • Builds a nested filtration at three cutoffs (2.0 Å, 3.0 Å, 4.0 Å) to capture covalent and non-covalent interactions.

  • Assigns Forman–Ricci curvature features to vertices, edges, and triangles at each filtration level.

Capability: The AI system can now:

  • Distinguish between polymers that differ only in repeating-unit boundary placement (translationally invariant).

  • Capture many-body interactions (3+ atoms simultaneously) that graph-based models miss.

  • Encode long-range non-covalent interactions (e.g., hydrogen bonding) that are critical for bulk properties like glass transition temperature.

The pretraining objective is a weighted sum of categorical cross-entropy (atom/bond) and binary cross-entropy (FG), with weights 2.0, 1.0, and 5.0 respectively.

The improved AI system can:

  1. Represent polymers as periodic, many-body, multiscale topological objects rather than finite molecular graphs.

  2. Predict a wide range of polymer properties (electronic, optical, physical, thermal) with state-of-the-art accuracy (lowest RMSE on 8/9 tasks, highest R2 on all 9).

  3. Quantify structure–property relationships with statistical rigor, including confidence intervals and significance tests.

  4. Provide reliable uncertainty estimates for individual predictions, enabling experimental prioritization.

  5. Generalize to novel polymers outside the training distribution, as demonstrated by experimental validation of newly synthesized materials.

  6. Guide polymer design by identifying functional group modifications that reliably enhance target properties (e.g., thermal stability).

Abstract

Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging. Most machine learning approaches represent polymers as molecular graphs of a single repeating unit, thereby missing both the periodicity of polymer chains and many-body interactions beyond pairwise bonds. We introduce Periodic-TDL, a deep learning framework built on periodic Vietoris-Rips complexes that capture many-body interactions across multiple spatial scales, followed by a hierarchical simplicial message-passing (HSMP) encoder that propagates information from long-range interactions to covalent bonds, yielding representations enriched by higher-order topological features. Periodic-TDL outperforms all state-of-the-art models across polymer property prediction tasks spanning electronic, optical, physical, and thermal targets. Furthermore, we quantitatively validate how ester-to-amide substitution and alpha-methylation enhance thermal stability. Using a computationally synthesized dataset of 48,208 structures-generated via systematic substitution of acrylate and acrylamide polymers-we observed a mean T g increase of about 55 C for ester-to-amide substitutions and about 14 C for backbone alpha-methylation across matched polymer pairs. To verify these predicted trends, we use our Periodic-TDL model to analyze six novel polymer pairs from independent experimental measurements, including three newly synthesized polymers previously unreported in the literature. The experimental data successfully confirmed the model's predictions. Ultimately, these findings demonstrate that Periodic-TDL captures the underlying physical effects of specific functional group modifications, rather than merely optimizing predictive performance on benchmark datasets.

Sources

Related papers