Bio papers — 2026-10-07

Today's focus is on improving how we predict R genes using homology, which is crucial because it helps us understand gene function in a more automated way. We tried HRPv2, an automated and enhanced method for full-length homology-based R gene prediction. This approach aims to find these genes by comparing sequences to known ones, which is a big step toward understanding how these genes work.

Another piece of work explored how causal analysis can reveal autonomy in models of biological systems. This method looks at complex biological models to see which parts are truly independent, giving us insight into the underlying mechanisms. Building on that idea, we looked at Phylo2Vec, which is a vector representation for binary trees used to encode phylogenetic relationships more effectively.

We also touched upon the relationship between discrete and continuous dynamics in biology, specifically when they align and when they diverge. This helps us decide which mathematical framework is best suited for describing biological processes at different scales. Furthermore, we investigated the identifiability of a linear dynamical reaction-diffusion system with directed interactions using only second-order statistics from a single snapshot.

Finally, we looked at how a linear fitness subspace within protein language models enables sample-efficient directed evolution. This work suggests that by understanding this subspace, we can guide evolutionary changes in proteins much more efficiently than traditional methods.

The work that truly matters is the development of mathematical invariant-enabled topological neural networks for predicting molecular and materials properties, which suggests a new way to model complex physical systems. This approach attempts to capture underlying structural symmetries in data, which could lead to more robust predictions than standard machine learning methods.

This network framework utilizes mathematical invariants to predict properties of molecules and materials, aiming for greater accuracy in material science applications. Furthermore, encoding level-three semi-directed phylogenetic networks using quarnets and quinnets provides a novel way to represent complex evolutionary relationships within biological data. This method offers a more sophisticated structure for analyzing how genetic information is organized across different species.

Another piece of work focuses on retinalysis-vascx, which is an explainable software toolbox designed for extracting retinal vascular biomarkers from images. This tool aims to make the extraction process transparent and understandable, which is crucial for clinical interpretation of eye health data. This contrasts with the database update IMPPAT 3.0, which refines a FAIR database of phytochemicals and formulations from Indian medicinal plants, providing a more accessible repository for ethnobotanical knowledge.

Today's papers

The papers

Important terms

R gene prediction
This is about automatically finding genes that cause disease by comparing their sequences to known ones, which helps us understand gene function in a more automated way.
Causal analysis
This method looks at complex biological models to figure out which parts are truly independent, giving us deep insight into the underlying mechanisms of biological systems.
Phylo2Vec
This is a new way to represent binary trees used for phylogenetic relationships. It helps encode evolutionary data in a more effective and sophisticated manner.
Topological neural networks
These are new mathematical models that use structural symmetries in data to predict properties of molecules and materials, aiming for more robust predictions than standard machine learning.