UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions

arXiv:2512.10515 · q-bio.BM · Submitted 2025-12-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: Today's paper: "UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions".

Marcus: The gist The UNAAGI model, a diffusion-based generative model that reconstructs residue identities from atomic-level structure using an E(3)-equivariant framework,

Ines: First, who's behind it and why it matters.

Title and authors: Ines: The paper's title, UNAAGI: Atom-Level Diffusion for Generating Non-Canonical Amino Acid Substitutions, points directly to its method, which is using diffusion at an atomic level to generate those different amino acid identities.

Marcus: It suggests that instead of just predicting a sequence of letters or tokens, we are generating the actual three-dimensional structure around the side chain atoms themselves.

Ines: That’s right, and the authors are Han Tang and Wouter Boomsma from DIKU in Copenhagen, Denmark. They set up this framework using an E(three)-equivariant setup for this atomic-level generation <ref:2512.10515#pg1,using an E(3)-equivariant>.

Marcus: So what's the big implication of that E(three)-equivariant framework <ref:2512.10515#pg1>? It means the model respects physical rules like rotation and translation when it’s learning about the molecular structure.

Yuki: That equivariance is key because protein structures are inherently three-dimensional, and modeling them that way should help it capture real spatial relationships between atoms.

The paper's summary: Ines: So what the UNAAGI paper actually summarizes is that they use a multi-modal diffusion approach where both the continuous atomic coordinates and discrete features are perturbed using different types of noise.

Marcus: They follow the standard training objective, which is minimizing that KL divergence per timestep while predicting the clean data, zero. It’s a standard diffusion setup adapted for this structure generation task.

Ines: The authors argue that by modeling side chains in full atomic detail instead of as discrete tokens, they can achieve some level of generalization from the natural amino acids to the non-canonical ones.

Marcus: They show that their method generates a histogram over a dynamic subset of chemical space relevant to the context, which includes both canonical and non-canonical amino acids.

Yuki: It’s an attempt to solve two big challenges at once: dealing with the vast chemical space of NCAAs and getting predictive power from data that only comes from natural amino acids.

The paper's improvements: Ines: The paper highlights a few specific improvements in their framework, like using an E(three)-equivariant Graph Neural Network to learn a score function over atom-level graphs <ref:2512.10515#pg1,using an E(3)-equivariant>.

Marcus: That means the network is designed to respect Euclidean transformations, which is important because molecular structures change shape when you rotate or move them in three dee space <ref:2512.10515#pg1>.

Ines: They also use a message passing mechanism where directional information between atoms is encoded as unit vectors, which helps build equivariant vector features during the learning process.

Marcus: And they introduce a virtual node strategy, inspired by DrugFlow, to handle the different side-chain sizes you get with natural versus non-canonical amino acids.

Yuki: That virtual node idea is smart because it lets them sample from a smoother distribution over identities across side chains of varying sizes without having to fix the atom count beforehand.

Conclusion: Ines: So, to wrap up the UNAAGI paper, they show that this atomic-level diffusion approach achieves meaningful correlations with experimental mutational effect data on ProteinGym.

Marcus: They found that UNAAGI shows a higher wild-type coverage rate compared to methods like PepINVENT, and importantly, it shows consistent performance across both canonical and non-canonical substitution benchmarks.

Yuki: For the wider context, this suggests that structure-based drug design principles might actually have a shared methodological foundation with protein engineering, opening up possibilities for unified training frameworks across these domains.

Ines: It’s a step forward because it’s the first diffusion-based approach to provide reliable predictive power on NCAA mutational effect benchmarks.

Marcus: The authors admit their limitation is that the samples tend to remain structurally close to the twenty natural amino acids, often interpolating between them rather than introducing radical variations <ref:2512.10515#pg2>.

Yuki: And another point they made is that in some benchmarks, UNAAGI only recovers a small subset of NCAAs like Nle, Nva, Abu, and tBu without reliably sampling other classes.

Department of Computer Science (DIKU) · University of Copenhagen

q-bio.BM

Submitted: 2025-12-11

Updated: 2026-10-07

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 82/100

The gist: The gist The UNAAGI model, a diffusion-based generative model that reconstructs residue identities from atomic-level structure using an E(3)-equivariant framework, achieves substantially improved

Key concepts

E(3)-equivariant framework
This mathematical framework ensures that the model respects geometric transformations like rotation and translation in 3D space. It means the model learns relationships between atoms that remain consistent regardless of how the protein structure is oriented, which is crucial for accurate structural modeling.
Atom-level side-chain generation
Instead of predicting a sequence of amino acids, UNAAGI generates the chemical identity and position of every atom in a residue. This detailed approach allows it to explore novel chemical combinations that go beyond the limitations of standard discrete sequence models, enabling the creation of non-canonical amino acid substitutions.
Diffusion-based generative model
The model uses a diffusion process, similar to how noise is gradually added and then reversed, to generate new structures or residue identities. This method allows it to sample from a complex distribution of possible atomic configurations, facilitating the generation of diverse and realistic non-canonical amino acid variants.

Terminology

Summary

The gist The UNAAGI model, a diffusion-based generative model that reconstructs residue identities from atomic-level structure using an E(3)-equivariant framework, achieves substantially improved performance on non-canonical amino acid substitutions compared to current state-of-the-art methods

Introduction and Motivation

Proposing beneficial amino acid substitutions, whether for mutational effect prediction or protein engineering, remains a central challenge in structural biology <ref:2512.10515#pg2>. Current approaches are constrained to designing sequences composed exclusively of natural amino acids (NAAs), while the larger set of non-canonical amino acids (NCAAs) remain largely inaccessible for current variant effect prediction methods <ref:2512.10515#pg3>. The paper explores a hypothesis that modeling amino acids generatively in full atomic detail, rather than as a set of discrete tokens, can obtain some level of generalization from the natural to the non-canonical amino acids <ref:2512.10515#pg3>. This is achieved by proposing Uncanonical Novel Amino Acid Generative Inference (UNAAGI), which reconstructs residue identities from atomic structure using an E(3)-equivariant framework <ref:2512.10515#pg3>.

Methodology: UNAAGI Framework

UNAAGI introduces a novel generative framework for residue-level sequence design based on atom-level side-chain generation via equivariant molecular diffusion <ref:2512.10515#pg3>. The model utilizes a multi-modal diffusion approach where both atomic coordinates (continuous) and atom-wise categorical features (discrete) are perturbed using Gaussian or categorical noise independently <ref:2512.10515#pg6>. The training objective follows the standard ELBO on the data log-likelihood, minimizing the per-timestep KL divergence which is optimized by predicting the clean data xˆ0 <ref:2512.10515#pg6>.

Key architectural components include:

  1. E(3)-equivariant Graph Neural Network: This network learns a score function over atom-level graphs, respecting Euclidean transformations such as rotation and translation through equivariance <ref:2512.10515#pg7>.

  2. Message Passing: Directional information between atoms is encoded as unit vectors xji,n = xj − xi / xj − xi, which are used to construct equivariant vector features during message passing <ref:2512.10515#pg7>.

  3. Virtual Node Strategy: To handle variable side-chain sizes across NAAs and NCAAs, UNAAGI employs a virtual node strategy inspired by DrugFlow, where nodes are assigned a special atom type NOATOM and can be removed post hoc during sampling <ref:2512.10515#pg7>.

Evaluation on Mutational Effect Prediction

The model is evaluated on experimentally benchmarked mutation effect datasets using Deep Mutational Scanning (DMS) data <ref:2512.10515#pg8>. Predictive power is quantified by comparing the learned distribution of side-chain identities against experimental DMS data, calculating the differential log-likelihood ∆ log L = − log P(mutant) + log P(wild-type) which is correlated with experimental ∆∆G values <ref:2512.10515#pg8>.

The results demonstrate that UNAAGI achieves meaningful correlations with experimental mutational effect across most assays on ProteinGym, exhibiting a higher wild-type coverage rate compared to baselines like PepINVENT <ref:2512.10515#pg9>. Crucially, UNAAGI shows consistent performance across both canonical and non-canonical substitution benchmarks, indicating it is the first diffusion-based approach to provide reliable predictive power on NCAA mutational effect benchmarks <ref:2512.10515#pg9>.

Analysis of Non-Canonical Amino Acids

The analysis of sampled non-canonical amino acids reveals two main phenomena: chemical diversity and limited structural deviation from natural amino acids <ref:2512.10515#pg10>. While UNAAGI generates a broad range of chemistries, the samples tend to remain structurally close to the 20 natural amino acids, often resembling existing residues or interpolating between them rather than introducing radical variations <ref:2512.10515#pg10>. Furthermore, in the Rogers et al. (2018) benchmark, UNAAGI consistently recovers only a subset of NCAAs included in the benchmark—specifically Nle, Nva, Abu, and tBu—while not reliably sampling other classes of NCAAs <ref:2512.10515#pg11>.

Conclusion and Future Directions

UNAAGI successfully bridges molecular generation and protein mutational effect prediction by modeling side chains atom by atom <ref:2512.10515#pg12>. The model's methodological foundation suggests a promising connection between structure-based drug design (SBDD) and variant effect prediction, opening the door for a unified training framework <ref:2512.10515#pg3>. Future work could integrate additional SBDD techniques, such as pharmacophore-based conditioning or flexible side-chain modeling, to further improve variant effect prediction <ref:2512.10515#pg12>. The current limitations include the scarcity of relevant NCAA data and the tendency to interpolate between canonical-like structures rather than sampling chemically distinct NCAAs <ref:2512.10515#pg12>.

--- Page 1 ---

Machine Learning for Structural Biology Workshop

UNAAGI: ATOM-LEVEL DIFFUSION FOR GENERATING NON-CANONICAL AMINO ACID SUBSTITUTIONS<ref:2512.10515#pg2>

--- Page 2 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg3>

--- Page 3 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg4>

--- Page 4 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg5>

--- Page 5 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg6>

--- Page 6 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg7>

--- Page 7 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg8>

--- Page 8 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg9>

--- Page 9 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg10>

--- Page 10 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg11>

--- Page 11 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg12>

--- Page 12 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg13>

--- Page 13 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg14>

--- Page 14 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg15>

--- Page 15 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg16>

--- Page 16 ---

Machine Learning for Structural Biology Workshop<ref:2512.10515#pg17>

--- Page 17 ---

Machine Learning for Structural Biology Workshop

--- Page 2512.10515 (Final)

Machine Learning for Structural Biology Workshop UNAAGI: ATOM-LEVEL DIFFUSION FOR GENERATING NON-CANONICAL AMINO ACID SUBSTITUTIONS<ref:2512.

Improvements for AI systems

  1. Bold header: E(3)-equivariant diffusion framework for atom-wise side-chain generation

UNAAGI's use of a E(3)-equivariant diffusion framework for atom-wise side-chain generation, offering a new approach to residue identity inference allows the system to propose plausible non-canonical substitutions by modeling the atomic coordinates of side chains directly, enabling exploration across both canonical and non-canonical amino acids within a unified paradigm.

  1. Bold header: Unified generative paradigm for protein engineering and structure-based drug design

The paper suggests that UNAAGI's methodology demonstrates a shared methodological foundation between protein engineering and structurebased drug design, opening the door for a unified training framework across these domains. This enables the system to leverage principles from both fields simultaneously for novel molecular generation tasks.

  1. Bold header: Predictive power on non-canonical amino acid (NCAA) substitutions

UNAAGI achieves consistent performance across both canonical and noncanonical substitution benchmarks, making it the first diffusion-based approach to provide reliable predictive power on NCAA mutational effect benchmarks. This allows the improved AI system to perform meaningful variant effect prediction for NCAAs, a previously inaccessible chemical space.

  1. Bold header: Variable side-chain modeling via virtual node strategy

The system employs a virtual node strategy inspired by DrugFlow (Schneuing et al., 2025) which allows for sampling from a smoother distribution over amino acid identities across side chains of varying sizes, addressing the limitation where in standard SBDD, the atom count must be fixed before diffusion.

  1. Bold header: Enhanced structural generalization for NCAAs

The model exhibits the ability to generalize to the subclass of non-canonical amino acids which are chemically proximal to the natural amino acids, suggesting it can predict substitutions for NCAAs that do not deviate drastically from canonical counterparts.

Abstract

A central challenge in structural biology is identifying beneficial amino acid substitutions, a problem that underpins both mutational effect prediction and protein engineering. Recent inverse-folding models, trained to reconstruct sequences from structure, have shown considerable promise for identifying functional mutations. However, current approaches are constrained to designing sequences composed exclusively of canonical amino acids (CAAs). Non-canonical amino acids (NCAAs) offer greater chemical diversity and are frequently used for protein engineering in vivo, yet they remain largely inaccessible to current variant effect prediction methods. To address this gap, we introduce UNAAGI, a diffusion-based generative model that reconstructs residue identities from atomic-level structure using an E(3)-equivariant framework. By modeling side chains in full atomic detail rather than as discrete tokens, UNAAGI enables the exploration of canonical and non-canonical amino acid substitutions within a unified generative paradigm. We evaluate UNAAGI zero-shot, without mutational-effect supervision or task-specific fine-tuning, on experimental mutational-effect benchmarks and demonstrate substantially improved performance on NCAA substitutions relative to existing methods.

Sources

Related papers