Retro3D: A 3D-aware Template-free Method for Enhancing Retrosynthesis via Molecular Conformer Information

arXiv:2501.12434 · cs.LG, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "ConfRetro: a 3D-aware template-free method for enhancing retrosynthesis via molecular conformer information".

Jane: The paper was written by Jiaxi Zhuang, Yu Zhang, Ying Qian and Aimin Zhou from East China Normal University and Shanghai Institute of Artificial Intelligence for Education and Shanghai Frontiers Science Center of Molecule Intelligent Syntheses.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Hey everyone, welcome back to the show! Today we're digging into a paper that's got a mouthful of a title — "Retrothree dee: A three dee-aware Template-free Method for Enhancing Retrosynthesis via Molecular Conformer Information." Jane, I gotta say, just reading that title made me want to grab a coffee and a chemistry textbook.

Jane: Ha, right? But honestly, the title tells you exactly what's going on. Retrosynthesis is this big problem in chemistry where you have a final molecule — say, a drug — and you need to figure out what starting materials you'd mix together to make it. It's like working backwards from a finished cake to the recipe.

Tom: And the "three dee" part is the twist here. Most previous methods treat molecules like flat strings of text or 2D graphs. But molecules are actually three-dimensional objects, right? Atoms are arranged in space, they have angles, distances, chirality — all that stuff matters for how reactions actually happen.

Jane: Exactly. And the authors, Zhuang and colleagues from East China Normal University, they're saying: look, if you ignore the three dee shape, you're going to make mistakes on complex molecules. Things like polychiral molecules — those have multiple chiral centers, like left-handed and right-handed versions — or heteroaromatic rings. Those are genuinely hard to get right without spatial awareness.

Tom: So they built a model called Retrothree dee that takes the molecule's three dee conformer — basically a snapshot of its shape in space — and feeds that into a Transformer, which is the same kind of architecture used in large language models. But here's the clever part: they don't just mash the three dee info in randomly. They have this "Atom-align Fusion" module that makes sure each atom's three dee position lines up with its corresponding token in the SMILES string.

Jane: Right, and that alignment is crucial. If you just concatenate features willy-nilly, the model might mix up which atom is which. But they're careful to keep the mapping consistent. And then they also have a "Distance-weighted Attention" mechanism that uses the actual three dee distances between atoms to guide how the model pays attention. So atoms that are close in space get more attention weight, which makes chemical sense.

Tom: I love that. It's like telling the model, "Hey, these two atoms are physically near each other, so pay attention to their relationship." That's a really intuitive way to inject chemistry knowledge without hard-coding reaction rules.

Jane: And the results back it up. On the USPTO-50K dataset, they beat all other template-free methods and even compete with template-based ones that rely on pre-defined reaction rules. We'll get into the numbers in a bit, but the big picture is: adding three dee information helps, and it helps a lot on tricky molecules.

Tom: So the title is a bit of a mouthful, but the idea is actually pretty elegant. Let's keep going — I want to hear more about how they actually pulled this off.

Summary: Tom: Alright, so we've got the gist of what Retrothree dee does. Jane, can you walk us through the actual method in a bit more detail? I know you've been reading the paper closely.

Jane: Sure thing. So the model has three main components. First, the Atom-align Fusion module — that's how they combine the SMILES sequence embedding with the three dee position embedding from a graph neural network called ComENet. ComENet is really good at capturing three dee geometric information like bond angles and distances. They pad the three dee embeddings so that non-atom tokens like parentheses or equals signs just get zeros, and then they use trainable weights to blend the two embeddings together.

Tom: So it's a learned fusion, not just a fixed concatenation. That's smart.

Jane: Exactly. Then there's the Distance-weighted Attention. This is where they take the three dee distance matrix — basically a table of how far apart every pair of atoms is — and they transform those distances using Gaussian basis functions. That gives them a high-dimensional representation of each distance. Then they use that to re-weight the attention scores in some of the attention heads. So the model has both normal attention heads, which learn sequence patterns, and spatial attention heads, which focus on three dee proximity.

Tom: And they refine those distance weights layer by layer, right? So the model is constantly updating its understanding of spatial relationships as it processes the molecule.

Jane: Yeah, they have a "Weight Refine" step that uses the updated hidden states to adjust the distance weights. It's like the model is learning which spatial relationships actually matter for the reaction. And then the third piece is SMILES Alignment — they use atom mapping between the product and reactants to create a guidance signal for the cross-attention in the decoder. That way, the model is encouraged to align tokens that correspond to the same atoms before and after the reaction.

Tom: That's a really nice use of supervision that's already in the data. The atom mapping is often available in reaction datasets, so they're leveraging that free signal.

Jane: Right. And the overall loss combines the standard language modeling loss with an R-Drop term for robustness and the alignment guidance loss. They trained it on USPTO-50K and also tested on the much larger USPTO-FULL dataset.

Tom: And the results? I saw some pretty impressive numbers in the abstract.

Jane: On USPTO-50K with reaction class unknown, they hit fifty-five point five percent Top-one accuracy. That's a solid jump over the previous best template-free method, Retroformer, which was around fifty-three point two percent. And on Top-ten they're at eighty-nine point one percent versus eighty-six point one percent for Retroformer. So it's a consistent improvement across the board.

Tom: And on the larger USPTO-FULL dataset, they also lead. That's important because it shows the method scales, which is often where template-based methods struggle.

Jane: Yeah, template-based methods need a predefined template database, and that doesn't scale well to huge datasets. Retrothree dee is template-free, so it learns the reaction rules implicitly from data. That's a big advantage.

Tom: So the summary is: three dee information helps, alignment matters, and the model is both accurate and scalable. What's not to love?

Jane: I'm curious about the limitations, though. Let's dig into what they improved and what's still left on the table.

Improvements: Tom: So we've covered the basics. But what does Retrothree dee actually improve over previous work, and why does it matter? Jane, you mentioned the numbers, but let's talk about the qualitative improvements too.

Jane: Great point. One of the most striking improvements is in validity. When you generate SMILES strings — those are the text representations of molecules — a lot of models produce invalid ones, especially at higher beam sizes. Retrothree dee achieves ninety-nine point eight percent Top-one validity on USPTO-50K, and even at Top-ten it's ninety-seven point zero percent. Compare that to Retroformer, which drops to ninety-six point seven percent at Top-ten. So the three dee information is helping the model generate chemically plausible structures.

Tom: That's huge. Invalid SMILES are useless in practice — you can't even feed them into a simulator. So a ninety-seven percent validity rate at Top-ten is a real practical win.

Jane: And then there's the round-trip accuracy. That's where they take the predicted reactants, run them through a forward reaction predictor, and see if they get back the original product. Retrothree dee hits ninety point seven percent Top-one round-trip accuracy, which is a massive jump from Retroformer's seventy-eight point nine percent. That means the predicted reactants are not just valid — they actually react to form the right product.

Tom: That's the real test, right? It's not enough to generate a molecule that looks plausible; it has to actually work in a reaction. And they're twelve percentage points better than the previous best. That's a game-changer for practical applications.

Jane: And the case studies are really compelling. They tested on polychiral molecules, heteroaromatic rings, fused or bridged ring systems, and complex molecules that combine multiple features. In each case, Retrothree dee got the correct prediction in the top one or two while a vanilla Transformer often failed to produce any valid SMILES in the top three.

Tom: So the three dee information is specifically helping on the hard cases — the ones where stereochemistry and spatial arrangement matter most. That's exactly where previous methods fell short.

Jane: And they also did an ablation study to show each component matters. Removing the SMILES alignment drops Top-one accuracy by about two percentage points. Removing the distance-weighted attention hurts too. And the full model is the best. So it's not just one trick — it's the combination of all three modules working together.

Tom: I also noticed they tested with different conformer generators and different three dee representation networks. The results were pretty stable, which suggests the method is robust to the quality of the three dee input.

Jane: Yeah, that's a nice robustness check. Even with a simpler three dee encoder like SchNet, they still get fifty-one percent Top-one accuracy, which beats many baselines. With ComENet, they get the full fifty-five point five percent. So the better the three dee representation, the better the results, but the method doesn't collapse if the three dee input is less sophisticated.

Tom: So what's the big takeaway for the field? It seems like three dee information is not just a nice-to-have — it's essential for accurate retrosynthesis, especially on complex molecules.

Jane: I think that's the key message. And it opens up a lot of future work — maybe integrating three dee info into semi-template methods, or using more accurate conformer generation from quantum chemistry calculations.

Tom: Let's bring in Lu and Meng — I'm curious what they think about the practical impact.

Lu: I think the biggest implication is for drug discovery. Retrosynthesis is a bottleneck in designing synthetic routes for new drug candidates. If you can predict reactants more accurately, especially for complex chiral molecules, you can save months of lab work. The twelve percent improvement in round-trip accuracy is not just a number — it means fewer failed experiments.

Meng: From an engineering standpoint, I'm impressed that they made it work end-to-end. The model is a Transformer with some clever modifications, and it trains in about thirty hours on a single RTX four thousand ninety. That's very accessible. You don't need a massive cluster to reproduce this.

Tom: So it's both scientifically novel and practically usable. That's a rare combination.

Jane: Definitely. And it sets a new bar for template-free retrosynthesis. Let's wrap up with some final thoughts.

Conclusion: Tom: Alright, we've had a great discussion about "Retrothree dee: A three dee-aware Template-free Method for Enhancing Retrosynthesis via Molecular Conformer Information." Jane, what's the one-sentence summary you'd give to a listener who just tuned in?

Jane: I'd say: Retrothree dee shows that adding three dee molecular shape information to a Transformer-based retrosynthesis model significantly improves accuracy, validity, and practical usefulness — especially for complex molecules that have been hard for previous methods.

Tom: And the key innovations — the Atom-align Fusion, the Distance-weighted Attention, and the SMILES Alignment — each contribute something meaningful. It's not a single trick; it's a well-engineered combination.

Jane: Right. And the results speak for themselves: state-of-the-art among template-free methods on both USPTO-50K and USPTO-FULL, with big jumps in round-trip accuracy that suggest the predictions are chemically sound.

Lu: I'd add that this could have a real impact on how chemists plan syntheses. If these predictions are reliable enough, they could be integrated into automated synthesis platforms, reducing trial and error in the lab.

Meng: And from a systems perspective, the fact that it trains on a single GPU and doesn't require a template database makes it much easier to adopt in real-world settings. That's a big deal for practical deployment.

Tom: So we're looking at a paper that's both academically strong and practically relevant. I'm excited to see where this line of research goes — maybe three dee-aware methods will become the standard for retrosynthesis.

Jane: Absolutely. And with that, we'll say goodbye to Retrothree dee and get ready to look at the next paper on our list. Thanks for listening, everyone!

Tom: See you next time!

Jiaxi Zhuang, Yu Zhang, Ying Qian, Aimin Zhou

East China Normal University · Shanghai Institute of Artificial Intelligence for Education · Shanghai Frontiers Science Center of Molecule Intelligent Syntheses

cs.LG, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

Code: https://github.com/Hanjun-Dai/GLN

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 72/100

Key concepts

Retrosynthesis
In chemistry, this is the process of working backward from a final molecule (like a drug) to determine what starting materials should be mixed together to create it. It is analogous to figuring out a recipe from a finished cake.
Molecular Conformer Information
Molecules are 3D objects, meaning their shape in space matters for chemical reactions. This information includes atoms' specific positions, angles, and distances, which the model uses to improve prediction accuracy.
Transformer
This is a type of neural network architecture (also used in large language models) that processes sequential data. In this context, it is adapted to take molecular information and predict reaction pathways.
Round-trip Accuracy
This measures the practical reliability of the prediction. It involves taking the predicted starting materials, running them through a forward reaction predictor, and checking if they successfully yield the original target product.

Terminology

Summary

Summary

Retrosynthesis is a fundamental problem in organic synthesis and drug discovery, focused on identifying a set of reactants capable of synthesizing a target product molecule. Existing approaches, while promising, often overlook 3D conformer information and molecular spatial structure, which limits their ability to generate reactants that comply with chemical rules, especially for complex molecules such as polychiral and heteroaromatic compounds. To address this, the paper proposes Retro3D, a transformer-based, template-free method that integrates molecular conformer information and spatial structure.

The paper identifies two main challenges for incorporating 3D conformer information into a template-free Transformer framework: (i) integrating 3D features while maintaining alignment between atom tokens and corresponding 3D representations, and (ii) constraining the model's receptive field based on spatial structure. To overcome these, Retro3D introduces two novel modules: Atom-align Fusion and Distance-weighted Attention.

The Atom-align Fusion module combines 3D positional information at the model input stage, ensuring alignment between atom tokens and their corresponding 3D representations. Specifically, it obtains a fusion embedding F 3D in R M times D by associating SMILES token embedding T in R M times D with 3D position embedding P 3D in R N times D, where M is the number of SMILES tokens and N is the number of atoms. The 3D position embedding is extracted using ComENet, a message-passing scheme that encodes radial distance, polar angle, azimuthal angle, and rotation angle within a 1-hop neighborhood. The fusion is computed as F 3D = lambda 1 times pad(P 3D) + lambda 2 times T, where lambda 1 and lambda 2 are trainable parameters that adaptively adjust the weights of the 3D position embedding and SMILES token embedding.

The Distance-weighted Attention mechanism guides self-attention by redistributing attention weights based on 3D distances between atoms. It constructs a 3D distance matrix D in R N times N from molecular conformers, then uses a Gaussian Basis Function to transform distances into a higher-dimensional space. The attention heads are divided into normal attention heads and spatial attention heads. Spatial attention heads use the 3D distance weight to limit the receptive field and emphasize chemically relevant atom pairs. The module also includes a Weight Refine step, which updates the 3D distance weight layer-by-layer using a fully connected layer that takes the concatenation of updated contexts of atom pairs.

The paper also introduces SMILES Alignment, which uses atom-mapping to construct a SMILES Alignment Map (SAM) between product and reactant tokens. This is incorporated into the training process via a guidance loss L SA that encourages alignment between cross-attention scores in the decoder's final layer and the SAM.

The overall loss function combines language modeling loss (including R-Drop loss for robustness) with the SMILES alignment guidance loss: L = L LM(CE) + alpha L LM(KL) + beta L SA, with alpha = 0.5 and beta = 1.0.

Experiments were conducted on the USPTO-50K and USPTO-FULL datasets. On USPTO-50K, Retro3D achieves state-of-the-art performance among template-free methods under both known and unknown reaction class conditions. For example, with reaction class unknown, Retro3D achieves Top-1 accuracy of 55.5%, Top-3 of 77.2%, Top-5 of 83.4%, and Top-10 of 89.1%, outperforming all template-free baselines. It also outperforms several template-based and semi-template methods. On USPTO-FULL, Retro3D achieves Top-1 accuracy of 50.8%, Top-3 of 72.9%, Top-5 of 78.0%, and Top-10 of 83.4%, again surpassing all baselines.

The paper also evaluates Top-k Validity and Top-k Round-Trip Accuracy. Retro3D achieves Top-1 validity of 99.8% and Top-10 validity of 97.0%, significantly higher than baselines like Graph2SMILES, RetroPrime, Retroformer, and NAG2G. For round-trip accuracy, Retro3D achieves Top-1 of 90.7%, Top-3 of 81.3%, Top-5 of 75.7%, and Top-10 of 69.0%, showing a remarkable improvement of 12% in Top-1 round-trip accuracy over the best baseline.

Ablation studies confirm the necessity of each module. Removing SMILES Alignment reduces Top-1 accuracy from 55.5% to 54.1%. Adding Atom-align Fusion improves Top-1 accuracy but slightly reduces Top-3 to Top-10 accuracy. Adding Distance-weighted Attention improves all metrics, and the combination of all three modules yields the best performance.

Case studies on molecules with intricate structures—polychiral, heteroaromatic, fused/bridged rings, and complex molecules—demonstrate that Retro3D can predict accurate and chemically plausible reactants, often at Top-1 or Top-2, while a vanilla Transformer without 3D information struggles to generate valid SMILES. For example, in a complex product case, Retro3D successfully deduces the molecular skeleton and achieves the ground truth at Top-1, while the vanilla Transformer fails to produce any valid SMILES within Top-3 predictions.

The paper concludes that Retro3D achieves a new state-of-the-art for template-free retrosynthesis methods and is highly competitive with template-based and semi-template methods. It also demonstrates potential for predicting synthesis routes for candidate drug molecules, as shown in multistep retrosynthesis examples for compounds like Febuxostat, Salmeterol, Nirmatrelvir, and Osimertinib.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:

  • Replace standard 1D SMILES-only encoding with a dual-stream encoder that processes both token sequences and 3D conformer coordinates simultaneously.

  • Implement the Atom-align Fusion module: instead of naive concatenation, use trainable weights (λ1, λ2) to blend 3D position embeddings (from ComENet) with SMILES token embeddings, ensuring token-to-atom alignment via zero-padding for non-atom tokens.

  • Modify the multi-head self-attention to split heads into two groups: normal heads (sequence context) and spatial heads (3D structure).

  • For spatial heads, compute a Gaussian-basis-expanded distance matrix between all atom pairs, transform it through a two-layer GELU network to produce a distance weight tensor, then apply it multiplicatively to the softmax attention scores before the final weighted sum.

  • Add a Weight Refine layer that updates the distance weights layer-by-layer using the concatenated hidden states of atom pairs.

  • During training, compute a ground-truth SMILES Alignment Map (SAM) from atom-mapped product/reactant pairs.

  • Add a cross-entropy loss term (β=1.0) that forces the final decoder layer's cross-attention scores to align with the SAM, strengthening the model's ability to track atom correspondence across the reaction.

  • Apply R-Drop loss (α=0.5) to the language modeling objective, ensuring robustness to dropout variations while the model learns 3D-aware representations.

  • Predict retrosynthesis reactants with higher accuracy, especially for molecules with complex 3D structures (polychiral, heteroaromatic, fused/bridged rings, and combinations thereof), achieving 55.5% Top-1 accuracy (vs. 50.4% for the best prior template-free method) on USPTO-50K without reaction class information.

  1. Spatially-Aware Generation: Generate reactant SMILES that respect 3D steric constraints, reducing invalid predictions (99.8% Top-1 validity vs. 99.4% for Graph2SMILES) and improving round-trip accuracy (90.7% Top-1 vs. 78.9% for Retroformer).

  2. Chirality Handling: Correctly predict stereochemistry in polychiral products by leveraging 3D conformer information, where 2D-only models fail (42.5% vs. 38.5% Top-1 accuracy on polychiral molecules).

  3. Reaction Center Identification: Use distance-weighted attention to focus on chemically relevant atom pairs (e.g., reactive bonds), improving predictions for heterocycle formation (47.0% Top-1 vs. typical 40-45% for baseline transformers).

  4. Scalability to Large Datasets: Handle USPTO-FULL (1M+ reactions) with 50.8% Top-1 accuracy, outperforming template-based methods that struggle with scale.

  5. Multistep Synthesis Planning: Recursively apply the model to generate complete synthetic pathways for drug-like molecules (e.g., Nirmatrelvir, Osimertinib), as validated in case studies.

  6. Robustness to Conformer Quality: Maintain performance even with lower-quality conformers (RDKit-generated vs. Schrödinger), showing only 0.3% accuracy drop, making it practical for real-world use.

  7. Interpretable Attention: Provide visualizable attention maps that correlate with 3D distance matrices, allowing chemists to verify the model's reasoning about molecular structure.

Sources

Related papers