Co-folding with a Soup of Representations

arXiv:2609.15552 · q-bio.BM · Submitted 2026-09-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: Today's paper: "Co-folding with a Soup of Representations".

Marcus: Co-folding models have rapidly advanced, but no single model consistently performs best across all biomolecular complexes,

Ines: First, who's behind it and why it matters.

Paper summary: Ines: So, summarizing what we've discussed about "Co-folding with a Soup of Representations," the paper argues that independently trained co-folding models encode complementary information that can be effectively transferred to improve structure prediction at inference time without retraining them.

Marcus: They specifically introduced this mechanism called SoupFold, which learns lightweight mappings between the representation spaces of different models, allowing it to aggregate these representations into a base model's representation before the final structure generation step.

Yuki: The significance lies in the fact that this aggregation successfully achieves state-of-the-art performance on both protein-protein and protein-ligand prediction tasks when tested against four existing co-folding models.

Ines: It really speaks to how different AI models, despite their similar foundational architectures, can still specialize in capturing different structural patterns that are mutually beneficial when combined.

Marcus: The implications for the field are that we might move toward a system where we don't rely on one monolithic model, but rather a curated soup of representations from several specialized models for the most accurate predictions.

Yuki: This suggests that understanding complex biological structures requires looking beyond the output of any single prediction tool to appreciate the broader landscape of structural possibilities encoded across many different AI approaches.

Conclusion: Ines: So, to wrap up this discussion on "Co-folding with a Soup of Representations," the core idea is that combining representations from different co-folding models helps predict protein structures better than any single model alone.

Marcus: Exactly, and I think what's really interesting is how they managed to do this without having to retrain all those massive base models, which saves a ton of computational power.

Yuki: From a population genetics standpoint, it’s fascinating that these distinct predictive capabilities—what we might call different "alleles" or structural patterns—can be aggregated into a more robust prediction for the entire species' protein landscape.

Ines: I'm curious about the authors themselves; what do they think is the big picture implication of finding this representation transfer mechanism?

Marcus: The authors suggest that it opens up a new way to utilize existing AI infrastructure, rather than having to build entirely new predictors from scratch for every complex.

Yuki: That implies a future where we can rapidly assess structural data across vast biological datasets by just querying the aggregated knowledge of these diverse models.

Ines: So, if we simplify it, the paper is basically showing us how to make a smarter prediction engine by letting different AI brains talk to each other during the final calculation.

Marcus: Right, and that means for genomics data scientists like myself, it's a powerful tool because we can use these combined predictions as a cross-validation check for our statistical models.

Yuki: It suggests that the underlying biological truth isn't locked in one single representation but is distributed across multiple, specialized views of the same complex.

Ines: That distributed view is what really intrigues me—does this mean we’re finally starting to see how structural information is encoded in a multi-layered way?

Marcus: It certainly does, and I think the next big thing will be seeing if we can adapt this representation mapping idea to other types of biological data where models aren't as well-established.

Yuki: That would really show how this principle applies beyond just protein folding, connecting it to how we understand evolutionary constraints on molecular shapes.

Hyosoon Jang, Taewon Kim, Sungsoo Ahn

KAIST

q-bio.BM

Submitted: 2026-09-14

Updated: 2026-09-28

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: Co-folding models have rapidly advanced, but no single model consistently performs best across all biomolecular complexes, raising the question of whether independently trained co-folding models

Key concepts

Co-folding Models
These are different AI models, like AlphaFold3 or ESMFold2, that all share a similar structure (Pairformer trunk). They predict protein structures but are trained on different recipes, meaning they often capture different structural patterns.
Representation Spaces
This refers to the internal mathematical spaces where each co-folding model stores the information about a complex. SoupFold learns how to translate or map these different spaces so that information from one model can be understood by another.
Teacher-to-Base Map
This is a learned function used during inference that translates the representation generated by a teacher model into the mathematical space of the base model. This allows the base model to incorporate useful structural insights from other models.
Complementary Information
This means that different models capture different aspects of a structure. Even if one model fails on a specific complex, its unique representation might hold information that is crucial for successful prediction when combined with another model's representation.

Terminology

Summary

Co-folding models have rapidly advanced, but no single model consistently performs best across all biomolecular complexes, raising the question of whether independently trained co-folding models encode complementary information that can be transferred across them. The introduction of SoupFold addresses this by learning simple mappings between the representation spaces of these models to improve structure prediction at inference time without retraining the original co-folding models.

The gist: SoupFold improves co-folding predictions by learning simple mappings between the representation spaces of co-folding models, allowing it to transfer and incorporate representations from other co-folding models to update the base model's representation before structure generation.

Introduction and Motivation

Co-folding models, such as AlphaFold3, Protenix, ESMFold2, and OpenDDE, share a similar architecture based on a Pairformer-based trunk that constructs pairwise representations of a complex. Despite this similarity in architecture, they differ significantly in their training recipes. The motivation for SoupFold stems from the observation that these models exhibit complementary behavior; for instance, one model might successfully predict one complex while another fails on the same complex. This suggests that different models may capture different structural patterns. The core question addressed is whether co-folding models encode complementary information in their representations, and can transferring this information improve another model? This relates to representation distillation, where a student matches a teacher's representation, but here the transfer occurs at inference time between independently trained co-folding models.

Methodology of SoupFold

SoupFold operates by designating one model as the base model (e.g., ESMFold2) for structure generation and treating the remaining co-folding models as teachers. The process involves learning lightweight mappings between their pair representation spaces using mean-squared error, specifically training a two-layer MLP with a hidden dimension of 1024 for each transfer network. At inference time, the teacher representations are mapped into the base model’s space via a teacher-to-base map and then incorporated into the base representation through an averaging scheme:

H = wbH(b) + P t wtft→b(H(t))

where wb and wt are coefficients of the base model and each teacher, respectively. In all experiments, these coefficients are set to 1.

Evaluation on FoldBench

SoupFold was evaluated on the FoldBench benchmark using four co-folding models: AlphaFold3, Protenix-v1, ESMFold2, and OpenDDE-preview. The results demonstrated that SoupFold achieves state-of-the-art performance on both protein-protein and protein-ligand structure prediction. Specifically, Table 1 shows that the +SoupFold model consistently outperforms all individual models across the metrics for Protein–Protein DockQ (≥ 0.23, ≥ 0.49, and ≥ 0.80) and Protein–Ligand success rates. The largest gains were observed with ESMFold2 for protein-protein prediction (a 6.93 percentage-point improvement at the high-quality threshold) and OpenDDE for protein-ligand prediction (from 55.38 to 63.44).

Analysis of Representation Utility

The study investigated whether the utility of a model's representation is determined by its standalone structure prediction performance. The findings indicate that this is not necessarily true; representations from models with lower standalone performance can still improve stronger models. For example, Table 2 shows that combinations involving Protenix and OpenDDE (A + O) improved the performance of ESMFold2 compared to using only one teacher. Furthermore, the transfer mechanism was shown to be selective: representation transfer mostly preserves the original prediction while shifting more interfaces toward higher-quality categories than lower-quality ones. This suggests that combining representations exploits information that individual models might miss, even when they fail on certain targets.

Recovery of Failed Predictions

A particularly significant finding is SoupFold's ability to recover predictions where all individual co-folding models fail. The authors illustrate cases in Figure 4 where all individual co-folding models fail to produce an acceptable structure, while SoupFold recovers acceptable predictions. This recovery cannot be explained by transferring the representation of a teacher that already predicts the correct structure; instead, combining representations from multiple models can recover useful information that is not sufficient for successful prediction in any individual model, providing strong evidence for complementary information encoded in their representations.

Future Directions

The paper suggests several avenues for future work. First, SoupFold may extend to other complex classes beyond protein-protein and protein-ligand complexes, with initial evidence showing improvement on antibody-antigen targets even when teacher models perform lower on that specific task. Second, using target-dependent teacher weights rather than a fixed weighting scheme could improve performance. Finally, extending representation transfer to other biomolecular models, such as protein or molecular foundation models (e.g.

Improvements for AI systems

Here are the specific improvements to AI systems based on the proposed SoupFold methodology:

  1. Enhance structure prediction accuracy for complex biomolecular assemblies (protein-protein and protein-ligand) by leveraging complementary information encoded in independently trained co-folding models.

  2. Implement a novel inference technique called SoupFold that dynamically transfers and incorporates learned representations from multiple state-of-the-art co-folding models (e.g., AlphaFold3, Protenix, ESMFold2, OpenDDE) into the representation space of a chosen base model (e.g., ESMFold2).

  3. Achieve state-of-the-art performance in protein structure prediction by combining representations at inference time via lightweight MLP-based transfer networks, avoiding the need for costly retraining of the original models.

  4. Improve predictive robustness by enabling the system to successfully predict structures for targets where all individual co-folding models fail, demonstrating that complementary information can provide structural guidance even when individual models are insufficient.

  5. Develop a mechanism to selectively improve predictions by identifying and transferring representations from teacher models that exhibit higher standalone performance or specific structural patterns useful for the base model, rather than relying solely on the teacher's final prediction accuracy.

The improved AI system (SoupFold) can perform the following:

  1. Predict stable three-dimensional structures of large, intricate biomolecular complexes with significantly higher accuracy (as evidenced by achieving state-of-the-art success rates on FoldBench tasks).

  2. Provide enhanced structural insights into protein-protein and protein-ligand interactions by effectively combining the diverse structural patterns captured by different model architectures.

  3. Serve as a powerful, low-overhead inference enhancement tool for existing co-folding models, allowing them to access knowledge distilled from other specialized models without requiring full model fine-tuning or expensive ensembling during the prediction phase.

  4. Act as a failover mechanism, providing high-confidence predictions for challenging or novel complex targets where traditional single-model predictors struggle.

Abstract

Co-folding models such as AlphaFold3, Protenix, ESMFold2, and OpenDDE have advanced rapidly, yet no single model consistently performs best across all biomolecular complexes. In this paper, we show that their pair representations encode complementary information that can be transferred across models to improve structure prediction. We introduce SoupFold, which combines pair representations from multiple co-folding models in a common representation space and generates structures from the combined representation. Importantly, SoupFold does not retrain the co-folding models and learns only simple mappings to transfer representations across models. We evaluate SoupFold on antibody-antigen, protein-protein, protein-ligand, molecular glue, GPCR, and oligomeric complex prediction using AlphaFold3, Protenix, ESMFold2, and OpenDDE. By combining representations across models, SoupFold improves over individual co-folding models across the considered benchmarks.

Related papers