Learning Topological Representations of Protein Structure and Dynamics

arXiv:2606.14737 · q-bio.BM, cs.LG, stat.ML · Submitted 2026-06-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: I'm Ines, and with me are Marcus and Yuki, guest researcher.

Marcus: Today's paper: "Learning Topological Representations of Protein Structure and Dynamics".

Ines: The gist Molecular dynamics (MD) simulations generate trajectories in a high-dimensional configuration space whose analysis critically depends on molecular descriptors, typically handcrafted observables or learned kinetic embeddings.

Marcus: First, who's behind it and why it matters.

Paper summary: Ines: So we're looking at this paper titled "Learning Topological Representations of Protein Structure and Dynamics" and the main idea is that molecular dynamics simulations give us trajectories, but analyzing that high-dimensional space depends on these molecular descriptors which are usually just handcrafted things or learned embeddings.

Marcus: Right, so the big question they tackle is how to design those descriptors that actually work well for everything. They propose using persistent homology as a general representation for MD and introduce something called the masked Flood complex, which is tailored specifically for proteins and aims to be computationally cheap while still capturing important structural information.

Ines: It seems like the thesis here is that these topologically informed summaries, when vectorized correctly, can give us rich data points for several different protein tasks simultaneously. They claim this shared representation space works for predicting protein classes, regressing physical observables at a frame level, and estimating Markov state models.

Yuki: From a population perspective, it’s interesting because they are focusing on making the underlying structure encoding specific to proteins right at the complex construction level instead of just applying the math after we have data. This integration of domain knowledge upstream is what they emphasize as a key shift in approach for protein structure representation.

Marcus: Exactly, and that leads into how they do it: they define this masked Flood complex by choosing landmarks as the C-alpha atoms and using masks to ensure that only atoms from different residues contribute to the filtration value, which keeps it focused on inter-residue structure while being efficient.

Ines: And then they take those persistence diagrams generated by this complex and vectorize them using a specific method based on exponential structure elements so that these summaries map into a coordinate system that is consistent across different molecular configurations.

Yuki: That consistency across configurations is important because it allows for population-level learning, which connects back to how we think about structural relationships in evolving species.

Marcus: They then use this vectorized representation to learn kinetic embeddings for Markov state models by employing the VAMPNet framework, optimizing a model jointly on a collection of trajectories to maximize the expected vamp-two score <ref:2606.14737#pg1>.

Ines: So, when they test this whole setup against protein class prediction, they say that PH-based representations substantially outperform the baselines and that their masked Flood PH yields the most consistent overall performance across classification and regression tasks.

Yuki: That consistency is what matters for a biologist; it means the summary isn't just good for one type of structural prediction but performs reliably across different structural domains.

Marcus: They also tested frame-level observable regression, predicting things like the radius of gyration or RMSD, and they found that both masked Flood PH and standard Flood PH performed best overall when looking at the average rank across those different tasks.

Paper summary: Ines: And for Markov state model estimation, which is crucial for understanding how a protein moves through its states, their mFlood PH ranks best among all the persistent homology based methods for metrics like vamp-two score and dynamical consistency tests <ref:2606.14737#pg1>.

Yuki: When you talk about MSM estimation and the stationary distribution agreement with Jensen–Shannon divergence, it means this representation is actually providing information about the dynamics of the system in a way that aligns with statistical expectations.

Marcus: On the generative side, they integrate these topological representations into a MarS-FM framework where a generative model is trained on MD trajectories using MSM-induced state transitions to sample frames for flow matching. They report that this MarS-FM mFlood model achieves better scores than the original MarS-FM and other topologically informed MSMs on the mdCATH dataset.

Ines: And they show promising transferability to fast folding proteins, where their model exhibits structural validation metrics like Bond RMSZ and Angle RMSZ that are comparable to MD reference trajectories. That’s a big deal for applying these concepts outside of the specific protein domains studied.

Yuki: It suggests that the way they’ve designed this complex doesn't just solve a problem for one set of data; it creates a representation language that can be used across different structural problems, which is what we see in evolution.

Marcus: So, to wrap up the summary of "Learning Topological Representations of Protein Structure and Dynamics," they show that masked Flood PH is the most consistent method among the persistent homology approaches when you look at classification, regression, and MSM estimation together.

Ines: The authors are essentially arguing that by integrating domain knowledge directly into how you build the simplicial complex for persistent homology, rather than just applying it generically to a point cloud of atoms, you create descriptors that are more informative about the underlying biology.

Yuki: The implication for the wider field is that we might be able to design principled ways to encode structural relations in a way that respects biological context, and this approach shows that biasing the persistent homology computation toward relevant inter-residue interactions works well.

Marcus: They also point out a limitation, though they state it plainly: the masked Flood complex isn't tied to one specific molecular system; it just means you can extend this idea to other domains where you can encode structural relations through masking.

Ines: It's about creating a robust and broadly informative summary for molecular dynamics that works across multiple downstream tasks, which is what they demonstrate with mFlood PH on the mdCATH dataset.

Yuki: The paper shows that incorporating this kind of domain knowledge directly into the simplicial complex provides a principled way to design descriptors for protein structure and dynamics, opening up new avenues for modeling these complex systems.

Marcus: So, in short, they developed the masked Flood complex to create a robust topological representation that performs consistently across classification, regression, and kinetic modeling tasks.

Conclusion: Ines: So, what we're hearing about today is that this paper, "Learning Topological Representations of Protein Structure and Dynamics," introduces a way to take the crazy high-dimensional data from molecular dynamics simulations and boil it down into something meaningful for biology.

Marcus: Yeah, basically they’ve got these complex mathematical tools called persistent homology and they’ve created a specific protein tool called the masked Flood complex that acts like a universal translator for that MD data.

Ines: It seems like the main point is that this representation isn't just some abstract math; it actually gives us useful summaries for predicting protein types, figuring out how the structure changes frame by frame, and even estimating how a protein moves through its different states.

Yuki: From a population view, this is interesting because they are using these structural summaries to connect different protein structures in a way that hints at how evolution might have shaped those shapes over time.

Marcus: Right, so the core of it is taking the raw atomic coordinates and turning them into these topological summaries, then vectorizing those summaries so we can actually learn stuff from them using AI models like VAMPNet.

Ines: And what they found is that this masked Flood approach performs consistently across all three major tasks they tested—classification, predicting physical measurements like the radius of gyration, and estimating the protein’s dynamics.

Yuki: That consistency is important because it means you get a reliable signal regardless of whether you're looking at classifying a whole family of proteins or just tracking a single one's movement.

Marcus: And they even showed that when they feed this into their generative model, MarS-FM, it gets better results than previous models on the mdCATH dataset and even shows promise for fast folding proteins.

Ines: So the big picture here is that by building domain knowledge directly into how you define these mathematical descriptors, you get a representation that's not just accurate for one thing but broadly informative across many different protein problems.

Yuki: And it’s not just about the accuracy of the numbers; it’s about building a language—a shared space—that allows us to connect structure to function in ways that reflect evolutionary history.

Marcus: It suggests that we can use these topological summaries to build much smarter AI models for understanding protein behavior, moving past just looking at individual atoms or single properties.

Dominik Geng, Florian Graf, Martin Uray, Roland Kwitt

University of Salzburg, Austria · Josef Ressel Centre for Intelligent and Secure Industrial Automation, University of Applied Sciences, Salzburg, Austria

q-bio.BM, cs.LG, stat.ML

Submitted: 2026-06-02

Updated: 2026-10-05

Comments: 36 pages, 6 figures

Code: https://github.com/google-deepmind/alphafold

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 90/100

The gist: The gist Molecular dynamics (MD) simulations generate trajectories in a high-dimensional configuration space whose analysis critically depends on molecular descriptors, typically handcrafted

Key concepts

Persistent Homology (PH)
PH is a mathematical tool used to analyze the shape of data by tracking topological features like holes and connected components as you vary a parameter. In this context, it helps capture the underlying structure of protein configurations, summarizing complex dynamics into simpler, stable topological features.
Masked Flood Complex (mFlood)
This is a specialized version of PH tailored for proteins. It modifies the standard Flood complex by focusing on specific structural elements like C$\alpha$ atoms and using masks to ensure that only interactions between different residues contribute to the topological summary, making it more relevant to protein structure.
Kinetic Embeddings
These are learned numerical representations derived from dynamic data, specifically used here for Markov State Model (MSM) estimation. By mapping molecular configurations into these embeddings, the model can learn how a protein moves between different states over time, aiding in understanding its dynamics.
VAMPNet Framework
This is a specific machine learning framework used to learn informative kinetic embeddings. It optimizes a model jointly on multiple MD trajectories to maximize a score called the vamp-2 score, which measures how well the learned representations capture the essential dynamic information.

Terminology

Summary

The gist Molecular dynamics (MD) simulations generate trajectories in a high-dimensional configuration space whose analysis critically depends on molecular descriptors, typically handcrafted observables or learned kinetic embeddings. The paper introduces persistent homology (PH) and a protein-tailored modification called the masked Flood complex as a general-purpose representation for MD, showing that these topologically informed descriptors provide information-rich summaries for tasks like protein class prediction, frame-level observable regression, and Markov state model (MSM) estimation in a single shared representation space.

How it works

  1. The core idea involves mapping atomic configurations to vectorized topological summaries via persistent homology, introducing an inductive bias about protein structure via an adapted simplicial complex construction that emphasizes structurally relevant interactions (Page 3).

  2. The masked Flood complex is a principled adaptation of a recent method for efficient persistent homology computation tailored to the characteristics of proteins (Page 3). It defines the masked flood complex mFloodr(X, L, Mσ) as a subcomplex of the original Flood complex from [22], where landmarks are chosen as the Cα atoms and masks ensure that only atoms of different residues contribute to the filtration value (Page 4).

  3. To utilize persistence diagrams in subsequent learning tasks, a vectorization based on exponential structure elements is adopted, which maps persistence diagrams (up to dimension 2) into a shared coordinate system consistent across molecular configurations (Page 3).

  4. This vectorized representation is then used to learn informative kinetic embeddings for MSM estimation by employing the VAMPNet framework, optimizing a model fθ jointly on a collection of trajectories to maximize the expected vamp-2 score (Page 5).

Representation Quality and Performance

The study evaluates three downstream tasks to probe representation quality:

** Task 1: Classification of top-level protein class:**

Vectorized persistence diagrams are used to assign each replica of an MD simulation of a protein domain to one of the four CATH classes by averaging them over time (Page 7). The results show that PH-based representations substantially outperform both baselines, with mFlood PH yielding the most consistent overall performance across classification, regression, and MSM estimation tasks (Page 9).

** Task 2: Regression of physical observables:**

The representations are tested on predicting common structural descriptors such as the radius of gyration (RG), secondary structure fraction (SSF), fraction of native contacts (FNC), and root mean-square deviation (RMSD) at the frame level (Page 8). Masked Flood PH and standard Flood PH yield the best overall performance as reflected by the lowest average rank across these tasks (Page 8).

** Task 3: MSM estimation:**

The quality of MSM estimation is measured using three complementary criteria: the vamp-2 score, dynamical consistency via Chapman–Kolmogorov (CK) tests, and agreement between the MSM-implied stationary distribution and the empirical distribution via Jensen–Shannon divergence (π-JS) (Page 8). mFlood PH ranks best among all PH-based methods for these metrics (Page 8).

Generative Modeling and Transferability

The topological representations are integrated into the MarS-FM framework of [31], where a generative model is trained on MD trajectories utilizing MSM-induced state transitions to sample (source, target) frames for flow matching (Page 9). The MarS-FM-mFlood model achieves better overall scores than the original MarS-FM model and other topologically informed MSMs on the mdCATH dataset (Page 9).

Furthermore, the study demonstrates promising transfer to fast folding proteins, where MarS-FM-mFlood exhibits structural validation metrics like Bond RMSZ and Angle RMSZ comparable to MD reference trajectories (Page 19).

Conclusion

Masked Flood PH delivers the most consistent performance among the considered PH-based representations across classification, regression, and MSM estimation tasks (Page 9). This representation provides a robust and broadly informative summary for molecular dynamics (Page 9). The use of topologically-informed MSMs enables sampling informative training pairs within the MarS-FM framework, leading to improved downstream generative modeling with better ensemble statistics on mdCATH and promising transfer to fast folding proteins (Page 9). This approach suggests that incorporating domain knowledge directly into the simplicial complex provides a principled way to design descriptors for protein structure and dynamics (Page 9). The masked Flood complex is not tied to molecular systems and naturally extends to other domains where structural relations can be encoded through masking (Page 9). This work shows that biased PH computation toward structurally relevant inter-residue interactions while retaining the GPU parallelism of Flood PH is effective (Page 9). The results indicate that mFlood PH exhibits the most consistent performance as seen in Table 4 (Page 9). This indicates that it provides a robust and broadly informative representation (Page 9). The paper concludes by showing that MSMs constructed from these representations enable sampling informative training pairs within the MarS-FM framework, leading to improved downstream generative modeling with better ensemble statistics on mdCATH and promising transfer to fast folding proteins (Page 9). The research was funded in whole or in part by the Austrian Science Fund (FWF) Grant-DOI 10.55776/DFH4791124 (Page 20). The paper is available as Preprint arXiv:2606.14737v1 [q-bio.BM] 2 Jun 2026 (Page 1). The paper's results on the mdCATH dataset show that PHbased descriptors are competitive across tasks, with masked Flood PH yielding the most consistent overall performance (Page 1). The authors acknowledge that masked Flood complexes are not tied to molecular systems and naturally extend to other domains where structural relations can be encoded through masking (Page 9). The paper's results on the mdCATH dataset show that PHbased descriptors are competitive across tasks, with masked Flood PH yielding the most consistent overall performance (Page 1). The authors acknowledge that masked Flood complexes are not tied to molecular systems and naturally extend to other domains where structural relations can be encoded through masking (Page 9). The paper's results on the mdCATH dataset show that PHbased descriptors are competitive across tasks, with masked Flood PH yielding the most consistent overall performance (Page 1). The authors acknowledge that masked Flood complexes are not tied to molecular systems and naturally extend to other domains where structural relations can be encoded through masking (Page 9). The paper's results on the mdCATH dataset show that PHbased descriptors are competitive across tasks, with masked Flood PH yielding the most consistent overall performance (Page 1). The authors acknowledge that masked Flood complexes are not tied to molecular systems and naturally extend to other domains where structural relations can be encoded through masking (Page 9). The paper's results on the mdCATH dataset show that PHbased descriptors are competitive across tasks, with masked Flood PH yielding the most consistent overall performance (Page 1). The authors acknowledge that masked Flood complexes are not tied to molecular systems and naturally extend to other domains where structural relations can be encoded through masking (Page 9). The paper's results on the mdCATH dataset show that PHbased descriptors are competitive across tasks, with masked Flood PH yielding the most consistent overall performance (Page 1). The authors acknowledge that masked Flood complexes are not tied to molecular systems and naturally extend to other domains where structural relations can be encoded through masking <ref:2606.

Improvements for AI systems

  1. Bold header: Masked Flood Complex for Structural Encoding

This complex provides a principled adaptation of a recent method for efficient persistent homology computation tailored to the characteristics of proteins, specifically by using Cα atoms as landmarks and masking higher-order simplices, leading to representations that are informative about the molecular system in a way that allows for resolving the system’s relevant processes.

  1. Bold header: Topologically-Informed Kinetic Embeddings

The system can learn informative kinetic embeddings for subsequent MSM estimation by using vectorized persistence diagrams as input into a VAMPNet framework, allowing it to learn a shared representation space across domains, which is crucial for capturing dynamic information across diverse protein ensembles.

  1. Bold header: Enhanced Generative Modeling with MarS-FM

The improved system can yield consistently better ensemble statistics on the mdCATH dataset when used within the MarS-FM framework, allowing it to sample (source, target) frames for flow matching guided by MSM-induced state transitions, leading to improved sampling of equilibrium distributions.

Abstract

Modern protein representation models support tasks such as enzyme design and drug discovery, but their reliance on static data such as sequence and native structure limits their ability to capture the conformational dynamics that drive protein function. We investigate whether persistent homology (PH) can provide descriptors shared across diverse proteins that retain global structure, fine-grained conformational variability, and kinetically relevant information without large-scale pretraining. We introduce the masked Flood complex, i.e., an adaptation of a recently proposed simplicial complex construction, that incorporates domain knowledge to emphasize inter-residue structure at low computational cost. We then use it to compute PH on molecular dynamics (MD) sampled structures, vectorize the persistence diagrams into a shared coordinate system, and probe the capacity of these representations in terms of the aforementioned aspects. To assess the amount of kinetic information, we learn low-dimensional embeddings from time-lagged observations and evaluate Markov state models (MSMs) estimated from them. Using these MSMs to guide training of the recent marsfm generative framework improves several ensemble statistics relative to the original model. After finetuning on lower-temperature MD data and adapting the sampling procedure, the resulting model also shows promising transfer to fast folding proteins.

Sources

Related papers