SPID: Distilled Protein Backbone Generation
cs.LG, cs.AI, stat.ML
Submitted: 2025-10-03
Updated: 2026-09-21
Comments: PyTorch implementation: https://github.com/LY-Xie/SiD_Protein
Code: https://github.com/LY-Xie/SiD_Protein
License: http://creativecommons.org/licenses/by/4.0/
The gist: Diffusion- and flow-based generative models have recently demonstrated strong performance in protein backbone generation tasks, offering unprecedented capabilities for de novo protein design.
Terminology
Abstract
Diffusion- and flow-based generative models have recently demonstrated strong performance in protein backbone generation tasks, offering unprecedented capabilities for de novo protein design. However, despite their generation quality, these models are constrained by slow sampling, often requiring hundreds of iterative steps. This computational bottleneck limits their practical utility in large-scale protein discovery, where thousands to millions of candidate structures are needed. To address this challenge, we explore the techniques of score distillation, which has shown great success in reducing the number of sampling steps in the vision domain while maintaining high generation quality. However, a straightforward adaptation of these methods results in unacceptably low designability. We introduce Score Protein identity Distillation (SPID), which resolves this incompatibility by combining few-step generation with inference-time noise scaling. SPID adapts the Score identity Distillation (SiD) framework to both diffusion- and flow-based models without requiring access to pretraining data. Applied to the Proteina flow-matching model, our 16-step generator achieves 94.4% designability, matching the 400-step teacher, while delivering more than a 20-fold reduction in effective backbone-generation time and maintaining comparable diversity and novelty. SPID generalizes across unconditional generation, fold-class conditional generation, and motif scaffolding, and extends to equivariant diffusion architectures, achieving significant reduction in generation time with comparable generation quality to the teacher in all tasks. The resulting reduction in inference cost could facilitate large-scale in silico protein design, thereby advancing diffusion-based models toward real-world protein engineering applications. The PyTorch implementation is available at https://github.com/LY-Xie/SiD Protein
Sources
- Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
- Protein Structure and Sequence Generation with Equivariant Denoising Diffusion Probabilistic Models
- SE(3)-Stochastic Flow Matching for Protein Backbone Generation
- A Survey on Pre-Trained Diffusion Model Distillations
- Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices
- Progressive Distillation for Fast Sampling of Diffusion Models
- Consistency Models
- Diffusion probabilistic modeling of protein backbones in 3D for the motif-scaffolding problem
- Fast protein backbone generation with SE(3) flow matching
- SE(3) diffusion model with application to protein backbone generation
- Few-Step Diffusion via Score identity Distillation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks