p2smi: A Python Toolkit for Peptide FASTA-to-SMILES Conversion and Molecular Property Analysis

arXiv:2505.00719 · q-bio.BM · Submitted 2025-04-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: I'm Ines, and with me are Marcus and Yuki, guest researcher.

Marcus: Today's paper: "p2smi: A Python Toolkit for Peptide FASTA-to-SMILES Conversion and Molecular Property Analysis".

Ines: p2smi is presented as a Python toolkit designed to bridge the gap between peptide sequence representations and chemical nomenclature, specifically focusing on converting peptide sequences into SMILES strings.

Marcus: First, who's behind it and why it matters.

Title and authors: Ines: So, to wrap up this discussion on "p2smi," we’re looking at how this paper summarizes what they actually delivered: a unified pipeline that takes a raw peptide sequence and reliably turns it into a standardized chemical string while simultaneously calculating key drug-like metrics like logP and TPSA for that resulting structure.

Marcus: I think their summary really highlights the practical, integrated nature of the toolkit; it’s not just one function, but a whole workflow where conversion and property assessment happen in tandem.

Yuki: From my viewpoint, their emphasis on handling structural complexity—like those various cyclization chemistries and unnatural amino acids—shows they are focused on capturing the real structural diversity we see across different biological species.

Ines: That makes sense; so they’re stressing that this toolkit manages both the sequence-to-string mapping and provides immediate chemical feasibility data, which is a lot of integrated functionality in one package. It seems like they're making the whole process much more cohesive than having to stitch together separate tools for different steps.

Marcus: Exactly; what’s compelling here for the genomics side is how this standardization helps tackle batch effects because we can generate sequences with defined modifications and then immediately check their physicochemical profiles consistently across different experimental groups. It streamlines that crucial data quality control step.

Yuki: I see the implication there as being huge for population genetics; if we can reliably convert sequence variants into chemical structures and assess their potential drug-likeness, it gives us a better framework to correlate specific genetic changes with physical properties in a way that was previously too cumbersome. It connects the genotype to the phenotype more directly.

Ines: So, they’re essentially arguing that by offering this comprehensive toolkit, they’re providing the necessary infrastructure for building high-throughput datasets that are chemically meaningful and statistically manageable for large-scale modeling. It's about making sure we aren't just getting sequences, but chemically realistic ones ready for analysis.

Marcus: That’s the statistical angle; having standardized property calculations means we can filter through vast sequence libraries much faster than before without needing a separate suite of tools for every single calculation. We cut down on redundant computational work significantly.

Yuki: I think it reinforces how this work connects the abstract genetic code to the physical reality of molecules, showing us how sequence variation translates directly into measurable chemical behavior in a way that's hard to track otherwise. It’s about seeing the molecular consequences of evolutionary pressures clearly reflected in the data.

Ines: It sounds like the main message is that p2smi lowers the barrier for using complex, modified peptides in large-scale computational modeling by making their input data perfectly compatible with standard cheminformatics tools. This makes it accessible to a much wider range of researchers.

Marcus: Precisely; it’s about making sure that when we feed these sequences into molecular dynamics or docking simulations, we’re giving the AI a clean, chemically accurate representation rather than something messy from a general converter. The quality of the input data directly dictates the quality of the simulation output.

Yuki: And looking at how they frame their future work, I think it shows they are already thinking about how this foundation can be expanded to incorporate more complex biological contexts beyond just basic drug-likeness checks. They aren't stopping at Lipinski’s rules; they're aiming for deeper structural understanding.

Ines: Right, and that leads us perfectly into what the authors suggest next: how these generated SMILES strings can actually be used to train and test generative models for designing novel peptide therapeutics. It shows they are already thinking about moving beyond just analyzing existing data toward active design.

The paper's summary: Ines: So we've talked about how p2smi functions as a toolkit for converting peptide FASTA sequences into SMILES strings and then analyzing their molecular properties, and now we need to talk about the actual improvements the authors propose for this toolkit, which I think is where things get really interesting for computational biology.

Marcus: I think their proposed improvements center around making the pipeline more modular so researchers can easily plug in different modification chemistries or property calculation methods without having to rewrite code from scratch every time.

Yuki: That modularity is critical because it allows us to test how different structural features—like adding a PEG chain versus a simple N-methylation—actually affect the resulting molecular properties in a controlled environment. It gives us experimental control over the structural variables we are testing.

Ines: So, they are suggesting an evolution of the toolkit where the conversion engine is decoupled from the modification and property prediction modules, which makes it much more flexible for varied experimental designs. This decoupling sounds like a big step toward making p2smi a true platform rather than just a fixed utility.

Marcus: I see that modularity directly impacting our cohort analysis; if we can rapidly swap out how a sequence is modified or analyzed, we can run massive comparative studies across different datasets much more efficiently without introducing new batch effects during the process. That efficiency gain for statistics is huge.

Yuki: That flexibility also speaks to the broader history of peptide study; it means we aren't locked into one single way of translating sequence data into chemical information, allowing us to explore structural space with greater freedom. It supports diverse hypotheses about how different genetic backgrounds manifest physically.

Ines: So, the core implication is that this toolkit isn't just a static converter anymore; it’s being designed as an adaptable framework for iterative experimental design and analysis in peptide research. It moves beyond one-off conversions to a continuous cycle of design, check, and refine.

Marcus: And from a data science view, it means the data generation part of our pipeline becomes much more scalable because we can automate the creation of highly diverse, chemically relevant training sets on demand. We're talking about on-demand library construction for AI training purposes.

Yuki: I think this points toward a future where we can use these tools to model not just single peptides but entire libraries of structurally diverse molecules that reflect different evolutionary pressures across species. It gives us the scope to study structural space more broadly.

Ines: So, the authors are looking at extending this capability beyond standard drug-like properties to incorporate more nuanced biological factors, which is a really exciting direction for computational biology. They aren't just checking if a peptide looks like a drug; they’re looking at deeper chemical relevance.

Marcus: And I think the next logical step they're pointing toward involves integrating this entire pipeline directly into generative AI frameworks, so the AI can propose modifications based on predicted outcomes rather than just random changes. This moves us from descriptive analysis to active, informed design.

Yuki: That would mean we move closer to an AI system that doesn't just predict a property, but actively designs a sequence with the intention of achieving that specific property. It’s about intelligent hypothesis generation based on chemical constraints.

Ines: It sounds like they’re moving from descriptive analysis—what is this peptide like?—to prescriptive design—what should we build next?—which is a significant step forward for understanding peptide-target interactions. This shifts the entire research paradigm from observation to creation.

The paper's improvements: Ines: So we've covered how p2smi functions as a toolkit for converting peptide FASTA sequences into SMILES strings and then analyzing their molecular properties, and now we need to wrap up what this all means for our field.

Marcus: Essentially, the paper lays out how this tool bridges the gap between raw genomic sequence data and chemically interpretable structures by providing reliable conversion and property assessment in one package.

Yuki: I think the real implication here is that we're getting a standardized language for discussing peptide structures derived from genetic variation across different populations, which helps us connect sequence differences to observable chemical outcomes.

Ines: That standardization is key; it means when we look at data from different cohorts, we can trust that the underlying molecular representations are being handled consistently by this toolkit.

Marcus: Exactly; it helps mitigate some of those batch effect issues because the conversion and property calculations are performed using a consistent methodology across all input sequences.

Yuki: And for population genetics, it gives us a standardized way to visualize how specific genetic markers translate into physical molecular properties that might affect species-specific function or stability.

Ines: It sounds like this work provides the necessary infrastructure for high-quality, chemically informed data generation, which is the foundation for building better predictive models in computational biology.

Marcus: That's right; it’s about ensuring that the input data into our AI models is as chemically accurate and statistically sound as possible before we start training anything.

Yuki: I think this kind of structured data handling is essential for understanding the long-term evolutionary trajectory of peptides in a species.

Ines: So, to sum up, p2smi delivers a practical, accessible way to translate complex peptide sequences into the chemical strings that drive property prediction and molecular modeling.

Marcus: It’s an incredibly useful resource for anyone working on high-throughput peptide screening or large-scale genomic data analysis because it makes the entire translation workflow much more streamlined.

Yuki: I think it really shows how fundamental tools like this, when applied thoughtfully, can provide deep insights into the physical constraints and evolutionary pressures acting on peptides.

Ines: We've seen that p2smi provides a robust foundation for sequence-to-string translation and property assessment for modified peptides.

Marcus: It’s a solid piece of software infrastructure that makes generating chemically meaningful datasets much more manageable for our large genomic cohorts.

Yuki: I think it underscores how vital it is to have these kinds of tools to connect the abstract genetic code with the physical reality of molecular chemistry in evolutionary studies.

Conclusion: Ines: So we've covered how p2smi functions as a toolkit for converting peptide FASTA sequences into SMILES strings and then analyzing their molecular properties, and now we need to wrap up what this all means for our field.

Marcus: Essentially, the paper lays out how this tool bridges the gap between raw genomic sequence data and chemically interpretable structures by providing reliable conversion and property assessment in one package.

Yuki: I think the real implication here is that we're getting a standardized language for discussing peptide structures derived from genetic variation across different populations, which helps us connect sequence differences to observable chemical outcomes.

Ines: That standardization is key; it means when we look at data from different cohorts, we can trust that the underlying molecular representations are being handled consistently by this toolkit.

Marcus: Exactly; it helps mitigate some of those batch effect issues because the conversion and property calculations are performed using a consistent methodology across all input sequences.

Yuki: And for population genetics, it gives us a standardized way to visualize how specific genetic markers translate into physical molecular properties that might affect species-specific function or stability.

Ines: It sounds like this work provides the necessary infrastructure for building high-quality, chemically informed data generation, which is the foundation for building better predictive models in computational biology.

Marcus: That's right; it's about ensuring that the input data into our AI models is as chemically accurate and statistically sound as possible before we start training anything.

Yuki: I think this kind of structured data handling is essential for understanding the long-term evolutionary trajectory of peptides in a species.

Ines: So, to sum up, p2smi delivers a practical, accessible way to translate complex peptide sequences into the chemical strings that drive property prediction and molecular modeling.

Marcus: It’s an incredibly useful resource for anyone working on high-throughput peptide screening or large-scale genomic data analysis because it makes the entire translation workflow much more streamlined.

Yuki: I think it really shows how fundamental tools like this, when applied thoughtfully, can provide deep insights into the physical constraints and evolutionary pressures acting on peptides.

Ines: We've seen that p2smi provides a robust foundation for sequence-to-string translation and property assessment for modified peptides.

Marcus: It’s a solid piece of software infrastructure that makes generating chemically meaningful datasets much more manageable for our large genomic cohorts.

Yuki: I think it underscores how vital it is to have these kinds of tools to connect the abstract genetic code with the physical reality of molecular chemistry in evolutionary studies.

Ines: We appreciate you all listening as we discuss p2smi: A Python Toolkit for Peptide FASTA-to-SMILES Conversion and Molecular Property Analysis, and look forward to hearing how these generated datasets are used in the next round of research.

Marcus: Indeed; it’s about making the process of translating complex peptide sequences into usable chemical strings manageable for large-scale computational studies, which is a necessary step before any meaningful property prediction can happen.

Yuki: It’s a useful development for understanding the structural diversity inherent in biological systems when we try to link sequence variation to actual molecular chemistry in an evolutionary context.

Department of Interdisciplinary Life Sciences, The University of Texas at Austin · Department of Integrative Biology, The University of Texas at Austin

q-bio.BM

Submitted: 2025-04-18

Updated: 2025-04-18

Comments: 4 pages

DOI: 10.21105/joss.08319)

Code: https://github.com/aaronfeller/p2smi

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: p2smi is presented as a Python toolkit designed to bridge the gap between peptide sequence representations and chemical nomenclature, specifically focusing on converting peptide sequences into SMILES

Key concepts

SMILES Conversion
This is the core function that translates a biological peptide sequence (like a string of amino acids) into a specific line-based chemical notation called SMILES. This notation allows computers to understand the 3D structure and chemical properties of the peptide, making it usable in cheminformatics workflows.
Noncanonical Amino Acids (NCAAs)
These are modified amino acids that are not part of the standard 20 building blocks found in natural proteins. p2smi supports over 100 different types of NCAAs, allowing researchers to model peptides with novel chemical properties that mimic or enhance natural peptide functions.
Molecular Properties Calculation
The toolkit computes essential chemical metrics for peptides, such as logP (a measure of lipophilicity), TPSA (topological polar surface area), and Lipinski's rule compliance. These calculations are crucial for predicting how a peptide might behave in biological systems or if it is likely to be a viable drug candidate.

Terminology

Summary

p2smi is presented as a Python toolkit designed to bridge the gap between peptide sequence representations and chemical nomenclature, specifically focusing on converting peptide sequences into SMILES strings. This tool addresses a critical need in computational modeling and cheminformatics by providing tailored functionalities for modified peptides, such as those incorporating noncanonical amino acids (NCAAs) and altered stereochemistry. By facilitating this conversion, p2smi reduces the overhead associated with generating accurate chemical strings for downstream analyses like molecular property calculations and drug-likeness assessments.

Statement of Need

The development of p2smi was driven by limitations in existing general bioinformatics toolkits, which often suffer from proprietary licensing or a lack of specific functionalities required for drug-like peptides. The primary motivation was the need to generate large-scale datasets of peptide SMILES strings for pretraining transformer-based models to understand SMILES notation. This necessity led to the creation of p2smi, which is built on core concepts from the CycloPs (Duffy et al., 2011) method for FASTA-to-SMILES conversion and has evolved into a stand-alone resource supporting peptide-focused machine learning pipelines.

Features

p2smi is designed to accommodate a wide range of complex peptide structures by leveraging external databases. Specifically, the toolkit accommodates over 100 unnatural amino acid residues by utilizing the database in SwissSidechain (Gfeller et al., 2012). Furthermore, it supports various structural complexities:

multiple cyclization chemistries, including disulfide bonds, head-to-tail, and sidechain cyclizations.

The toolkit also offers modification capabilities to explore altered peptide behaviors. For instance, it offers a SMILES modification tool, allowing users to apply N-methylation and PEGylation—modifications often used to influence peptide-drug stability and bioactivity. Finally, for early-stage drug evaluation, it computes key molecular properties such as logP, TPSA, molecular weight, and Lipinski’s rule compliance.

Command-Line Tools

p2smi is implemented as a pip-installable package offering five primary command-line tools to facilitate different stages of peptide analysis and modification. These tools are:

  1. The generate-peptides tool, which enables the generation of random peptide sequences based on user-defined constraints and modifications, allowing for the creation of diverse peptide libraries for computational studies.

  2. The fasta2smi tool, which performs the core conversion task: Converts peptide sequences from FASTA format into SMILES notation, facilitating integration with cheminformatics workflows.

  3. The modify-smiles tool, which allows users to apply chemical changes: Applies specific chemical modifications, such as N2 methylation and PEGylation, to existing SMILES strings.

  4. The smiles-props tool, which handles property calculation: Computes molecular properties—including logP, topological polar surface area (TPSA), molecular formula, and evaluates compliance with Lipinski’s rules.

  5. The synthesis-check tool, which assesses practical synthesis: Evaluates the synthetic feasibility of peptides based on defined synthesis rules, aiding researchers in determining the practicality of synthesizing specific peptide sequences.

State of the Field

While existing tools like pyPept and PepFuNN have made significant contributions to peptide informatics—with pyPept focusing on generating 2D/3D representations and PepFuNN on structure–activity relationship analyses—they lack a dedicated, direct conversion capability. The paper positions p2smi as complementary to these efforts, stating that pyPept and PepFuNN focus on structural representation, analysis, and structure–activity relationship studies of peptides, complementing the sequenceto-SMILES conversion capabilities provided by p2smi. This distinction highlights p2smi's unique role in providing the foundational sequence-to-string translation necessary for large-scale data generation.

Code Availability

The toolkit is made accessible to the research community through open access. p2smi is available as a pip-installable package on PyPI at https://pypi.org/project/p2smi, and the source code, including documentation and example notebooks, is openly available on GitHub at https://github.com/aaronfeller/p2smi. The work was supported by NIH grant 1R01 AI148419 and the Blumberg Centennial Professorship in Molecular Evolution at The University of Texas at Austin.

References

**(The paper lists references including ChemAxon, Duffy et al., Feller & Wilke, Gfeller et al., Landrum, O’Boyle et al., Ochoa et al., and OpenEye.

Improvements for AI systems

Here are specific improvements to AI systems based on the p2smi toolkit, along with what these improved systems can achieve:

  1. Improved Peptide Sequence Generation for Large-Scale Data Pretraining:

Identify and generate massive, diverse datasets of peptide SMILES strings incorporating Noncanonical Amino Acids (NCAAs), various backbone modifications, and complex cyclizations directly via the p2smi toolkit's CLI. This capability allows AI models (like transformer-based language models) to be pretrained on significantly larger, chemically realistic peptide sequence-to-SMILES mappings than current methods allow.

  1. Enhanced Peptide Property Prediction for Drug Design:

Integrate the p2smi pipeline into machine learning workflows to predict critical molecular properties (LogP, TPSA, Molecular Weight) and Lipinski's rule compliance directly from a raw peptide sequence or its SMILES representation. This allows AI systems to rapidly filter vast chemical spaces, prioritizing peptides with predicted favorable drug-like characteristics before costly synthesis or experimental validation.

  1. Automated Chemical Modification for Library Exploration:

Implement the p2smi modification tool (e.g., N-methylation, PEGylation) as a module within generative AI frameworks. This enables the AI to propose and test chemically relevant analogs of existing peptides by automatically applying these modifications, allowing researchers to explore structure-activity relationships (SAR) efficiently in silico without manual intervention in cheminformatics pipelines.

  1. Synthetic Feasibility Prediction:

Develop an AI system that leverages the p2smi synthesis-check function to predict the practical synthetic route for a given peptide sequence. This capability acts as a crucial gatekeeper, preventing the generation of computationally promising but synthetically intractable peptide designs, thereby reducing experimental waste and accelerating lead optimization cycles.

  1. Bridging Sequence-to-Structure Gap for Molecular Modeling:

Utilize p2smi's SMILES conversion function within molecular simulation pipelines (e.g., docking or molecular dynamics). This allows AI systems to seamlessly transition from a high-level peptide sequence input to the standardized chemical representation required by established cheminformatics tools, enabling the application of complex physical simulations (like predicting peptide diffusion across membranes) directly on generated libraries.

Abstract

Converting peptide sequences into useful representations for downstream analysis is a common step in computational modeling and cheminformatics. Furthermore, peptide drugs (e.g., Semaglutide, Degarelix) often take advantage of the diverse chemistries found in noncanonical amino acids (NCAAs), altered stereochemistry, and backbone modifications. Despite there being several chemoinformatics toolkits, none are tailored to the task of converting a modified peptide from an amino acid representation to the chemical string nomenclature Simplified Molecular-Input Line-Entry System (SMILES), often used in chemical modeling. Here we present p2smi, a Python toolkit with CLI, designed to facilitate the conversion of peptide sequences into chemical SMILES strings. By supporting both cyclic and linear peptides, including those with NCAAs, p2smi enables researchers to generate accurate SMILES strings for drug-like peptides, reducing the overhead for computational modeling and cheminformatics analyses. The toolkit also offers functionalities for chemical modification, synthesis feasibility evaluation, and calculation of molecular properties such as hydrophobicity, topological polar surface area, molecular weight, and adherence to Lipinski's rules for drug-likeness.

Related papers