UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions

summary

Video file (mp4)

The gist

The gist The UNAAGI model, a diffusion-based generative model that reconstructs residue identities from atomic-level structure using an E(3)-equivariant framework, achieves substantially improved

In short

UNAAGI is a generative model that reconstructs residue identities from atomic structures using an E(3)-equivariant diffusion framework. It aims to improve prediction of effects from non-canonical amino acid substitutions by modeling amino acids at the atom level rather than as discrete tokens. The model shows promising predictive power across both natural and non-canonical substitution benchmarks, suggesting a link between structure generation and variant effect prediction.

Key concepts

E(3)-equivariant framework
This mathematical framework ensures that the model respects geometric transformations like rotation and translation in 3D space. It means the model learns relationships between atoms that remain consistent regardless of how the protein structure is oriented, which is crucial for accurate structural modeling.
Atom-level side-chain generation
Instead of predicting a sequence of amino acids, UNAAGI generates the chemical identity and position of every atom in a residue. This detailed approach allows it to explore novel chemical combinations that go beyond the limitations of standard discrete sequence models, enabling the creation of non-canonical amino acid substitutions.
Diffusion-based generative model
The model uses a diffusion process, similar to how noise is gradually added and then reversed, to generate new structures or residue identities. This method allows it to sample from a complex distribution of possible atomic configurations, facilitating the generation of diverse and realistic non-canonical amino acid variants.

Terminology used across episodes

This episode discusses

The paper

UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions · Read on arXiv

Department of Computer Science (DIKU) · University of Copenhagen

A central challenge in structural biology is identifying beneficial amino acid substitutions, a problem that underpins both mutational effect prediction and protein engineering. Recent inverse-folding models, trained to reconstruct sequences from structure, have shown considerable promise for identifying functional mutations. However, current approaches are constrained to designing sequences composed exclusively of canonical amino acids (CAAs). Non-canonical amino acids (NCAAs) offer greater chemical diversity and are frequently used for protein engineering in vivo, yet they remain largely inaccessible to current variant effect prediction methods. To address this gap, we introduce UNAAGI, a diffusion-based generative model that reconstructs residue identities from atomic-level structure using an E(3)-equivariant framework. By modeling side chains in full atomic detail rather than as discrete tokens, UNAAGI enables the exploration of canonical and non-canonical amino acid substitutions within a unified generative paradigm. We evaluate UNAAGI zero-shot, without mutational-effect supervision or task-specific fine-tuning, on experimental mutational-effect benchmarks and demonstrate substantially improved performance on NCAA substitutions relative to existing methods.

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: Today's paper: "UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions".

Marcus: The gist The UNAAGI model, a diffusion-based generative model that reconstructs residue identities from atomic-level structure using an E(3)-equivariant framework,

Ines: First, who's behind it and why it matters.

Title and authors: Ines: The paper's title, UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions, points directly to its method, which is using diffusion at an atomic level to generate those different amino acid identities.

Marcus: It suggests that instead of just predicting a sequence of letters or tokens, we are generating the actual three-dimensional structure around the side chain atoms themselves.

Ines: That’s right, and the authors are Han Tang and Wouter Boomsma from DIKU in Copenhagen, Denmark. They set up this framework using an E(three)-equivariant setup for this atomic-level generation <ref:2512.10515#pg1,using an E(3)-equivariant>.

Marcus: So what's the big implication of that E(three)-equivariant framework <ref:2512.10515#pg1>? It means the model respects physical rules like rotation and translation when it’s learning about the molecular structure.

Yuki: That equivariance is key because protein structures are inherently three-dimensional, and modeling them that way should help it capture real spatial relationships between atoms.

The paper's summary: Ines: So what the UNAAGI paper actually summarizes is that they use a multi-modal diffusion approach where both the continuous atomic coordinates and discrete features are perturbed using different types of noise.

Marcus: They follow the standard training objective, which is minimizing that KL divergence per timestep while predicting the clean data, zero. It’s a standard diffusion setup adapted for this structure generation task.

Ines: The authors argue that by modeling side chains in full atomic detail instead of as discrete tokens, they can achieve some level of generalization from the natural amino acids to the non-canonical ones.

Marcus: They show that their method generates a histogram over a dynamic subset of chemical space relevant to the context, which includes both canonical and non-canonical amino acids.

Yuki: It’s an attempt to solve two big challenges at once: dealing with the vast chemical space of NCAAs and getting predictive power from data that only comes from natural amino acids.

The paper's improvements: Ines: The paper highlights a few specific improvements in their framework, like using an E(three)-equivariant Graph Neural Network to learn a score function over atom-level graphs <ref:2512.10515#pg1,using an E(3)-equivariant>.

Marcus: That means the network is designed to respect Euclidean transformations, which is important because molecular structures change shape when you rotate or move them in three dee space <ref:2512.10515#pg1>.

Ines: They also use a message passing mechanism where directional information between atoms is encoded as unit vectors, which helps build equivariant vector features during the learning process.

Marcus: And they introduce a virtual node strategy, inspired by DrugFlow, to handle the different side-chain sizes you get with natural versus non-canonical amino acids.

Yuki: That virtual node idea is smart because it lets them sample from a smoother distribution over identities across side chains of varying sizes without having to fix the atom count beforehand.

Conclusion: Ines: So, to wrap up the UNAAGI paper, they show that this atomic-level diffusion approach achieves meaningful correlations with experimental mutational effect data on ProteinGym.

Marcus: They found that UNAAGI shows a higher wild-type coverage rate compared to methods like PepINVENT, and importantly, it shows consistent performance across both canonical and non-canonical substitution benchmarks.

Yuki: For the wider context, this suggests that structure-based drug design principles might actually have a shared methodological foundation with protein engineering, opening up possibilities for unified training frameworks across these domains.

Ines: It’s a step forward because it’s the first diffusion-based approach to provide reliable predictive power on NCAA mutational effect benchmarks.

Marcus: The authors admit their limitation is that the samples tend to remain structurally close to the twenty natural amino acids, often interpolating between them rather than introducing radical variations <ref:2512.10515#pg2>.

Yuki: And another point they made is that in some benchmarks, UNAAGI only recovers a small subset of NCAAs like Nle, Nva, Abu, and tBu without reliably sampling other classes.

More episodes

← Home