A likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.
Ines: Today's paper: "A likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks".
Marcus: A likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks introduces an extension to existing Biologically-Informed Neural Networks (BINNs) that allows for the direct discovery…
Ines: First, who's behind it and why it matters.
Title and authors: Ines: So, moving on to what the paper actually says, they outline a framework that modifies the standard Biologically-Informed Neural Network setup to include this learnable noise component. Basically, they are using a likelihood-based approach where the model has to satisfy both the observed data and an ordinary differential equation simultaneously.
Marcus: I’m looking at how they formulate this; it seems like they replace that fixed Gaussian noise assumption with something more flexible, specifically modeling observations as u o i = u(t i) + sigma(u(t i)) epsilon, where epsilon is just standard normal noise.
Ines: That structure is key because it lets them define a power-law for that noise magnitude, sigma(u) = sigma0u alpha, and they treat the parameters of that scaling law, specifically alpha and sigma0, as learnable quantities directly from the data.
Marcus: That’s interesting because it means the AI isn't just fitting a single noise variance; it's discovering if the noise scales linearly or with some other power relationship based on density.
Yuki: From an evolutionary perspective, this suggests that different population states might be subject to fundamentally different levels of inherent uncertainty, which has implications for how we model species resilience.
Ines: They show that by incorporating this heteroscedastic noise term into the total loss function—which combines the data likelihood loss with the ODE residual and a biological constraint loss—they can actually recover both the underlying dynamics and that specific noise scaling law.
Marcus: The total loss function structure is quite neat because it balances three different types of constraints, giving equal weight to fitting the data, satisfying the growth equation, and keeping things biologically sensible like ensuring densities stay positive.
Ines: That balancing act is what makes this framework powerful; it forces the AI to find a solution that respects all these different aspects of reality at once.
Marcus: It sounds like they are proving that you don't have to pre-specify the noise form; the system can infer it from how noisy its data is when you look at different parts of the population.
Yuki: That inference capability, especially regarding density dependence, could provide crucial context for interpreting long-term ecological time series where density changes are continuous.
Ines: So, in short, they’re using a likelihood framework to simultaneously infer the growth dynamics and a data-driven noise model that scales with population size.
Marcus: And the results they show on synthetic logistic and Richards’ models are pretty compelling because they actually managed to recover different noise behaviors depending on the underlying model tested.
Yuki: Recovering different scaling laws like constant variance in one case versus linear variance in another shows the method isn't just a general fit but can identify specific biological processes.
Ines: That ability to distinguish between different noise regimes is what really elevates this work above existing methods that only assume simple Gaussian noise.
Marcus: It’s a solid demonstration that the framework works across different growth models, which suggests it has some general applicability beyond just one specific type of population growth curve.
Yuki: If this holds up on real biological data, it means we can start making much more nuanced predictions about population fluctuations in complex ecosystems.
Ines: So, we’re looking at a method that moves from assuming noise is a nuisance to treating it as a mechanistic quantity that needs to be learned alongside the system dynamics.
Marcus: It certainly shifts the focus of what we consider "error" in these models from just data misfit to understanding the physical process generating that error.
Yuki: That shift in perspective is exactly where I think we need to be when we’re trying to understand how species cope with environmental stress.
The paper's summary: Ines: Now, let's talk about what improvements this framework actually suggests over the methods that came before it. The main improvement is moving away from the fixed homoscedastic noise assumption, which was a major limitation in previous Biologically-Informed Neural Networks.
Marcus: That’s right; previous approaches often just assumed constant noise variance, which means they couldn't account for situations where variability naturally increases or decreases as the population density changes.
Ines: This paper proposes a learnable noise model sigma(u) = sigma0u alpha, which allows the framework to discover that scaling law directly from the observed data, meaning alpha and sigma0 become parameters inferred from the data itself.
Marcus: That’s significant because it means the AI isn't just trying to fit a single error term; it can capture complex, density-dependent variability structures present in real biological systems.
Yuki: From a species perspective, this is important because we know that uncertainty in population counts or growth rates often increases near carrying capacity or when populations are very small.
Ines: Exactly; the framework allows for an accurate and interpretable quantification of heteroscedastic uncertainty, which goes far beyond just giving you a simple Gaussian error bound on the prediction.
Marcus: If we can calibrate this uncertainty across different regimes—additive noise, multiplicative noise, or something in between—then our predictions will be much more useful for risk assessment in biological systems.
Ines: Furthermore, the paper shows that when you train with this loss function instead of just standard RMSE loss, you get a lower mechanistic error when trying to reconstruct the underlying growth law itself.
Marcus: That’s a crucial point; it suggests that by training on the noise structure alongside the dynamics, we get a more robust inference about the actual biological process driving those dynamics.
Yuki: If this leads to better reconstruction of the growth laws, it directly impacts our ability to understand how species respond to environmental pressures over evolutionary timescales.
Ines: So, in essence, they’ve improved the system by making it smarter about what kind of noise it's seeing and how that relates to the population state.
Marcus: It seems like they’ve moved from just getting a prediction to understanding the uncertainty surrounding that prediction in a much more detailed way.
Yuki: This capability to model density-dependent noise structure is exactly what we need when analyzing complex, real-world ecological data where variability isn't uniform across the sample.
The paper's improvements: Ines: So, to wrap up on this paper, the authors introduce a likelihood-based framework for simultaneously learning both the growth dynamics and a density-dependent noise model using biologically-informed neural networks. The main implication is that we can now discover how uncertainty arises from the system itself rather than treating it as just random error.
Marcus: And for us in data science, it means we have a way to get calibrated uncertainty intervals that actually reflect the known noise regimes, which is much more useful than generic error bars on top of a model.
Yuki: I think the most important implication for population biology is gaining mechanistic insight into how uncertainty scales with population density across different growth models, which helps us understand species resilience better.
Ines: Indeed, and by successfully recovering different noise scaling laws from synthetic data—like constant variance versus linear scaling—the framework offers a way to identify the specific noise process at play.
Marcus: It’s a significant step forward because it demonstrates that incorporating this learnable noise model improves both the accuracy of the inferred dynamics and provides a more robust estimate of those dynamics.
Yuki: This work, "A likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks," gives us a tool to analyze real experimental data where we can infer unknown growth laws alongside their associated variability structure.
Ines: That’s the core contribution; treating noise as a mechanistic quantity instead of just ignoring it, which is really valuable for computational biology.
Marcus: It’s a solid piece of work that shows how combining likelihood theory with BINN architectures can yield more informative and robust results when dealing with noisy biological time-series data.
Yuki: We're genuinely excited about this paper because it moves us closer to modeling complex ecological processes with greater accuracy and mechanistic understanding.
Conclusion: Ines: So, we've talked about how this paper introduces a likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks, and now we're wrapping up with some final thoughts on its impact.
Marcus: It’s really clear that the authors managed to move beyond just fitting the data by actually inferring the underlying noise structure directly from the observations, which is a major statistical win for analyzing complex biological cohorts.
Yuki: From a population genetics standpoint, having a method that can distinguish between different noise regimes in growth models means we can finally start building more nuanced models of species evolution under varying environmental stresses.
Ines: Exactly; the ability to recover the specific scaling law of observational variability is what makes this approach so powerful for understanding how uncertainty is structured in these systems.
Marcus: And that improved mechanistic error when training on noise structure instead of just raw data misfit suggests a more reliable inference about the actual biological process governing those time-series.
Yuki: I think it opens up avenues for testing hypotheses about species resilience in real-world data where we don't have perfect replicated measurements, which is a huge practical benefit.
Ines: So, to recap, this paper shows a framework that simultaneously infers growth laws and density-dependent noise structure using BINN techniques.
Marcus: That's right; it’s about treating noise as a learnable part of the mechanism rather than just an unwanted nuisance in the data fitting process.
Yuki: It definitely sets a new direction for how we might analyze ecological time series, providing a way to quantify uncertainty that respects density-dependent variability.
Ines: This study on "A likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks" really shows the power of linking observation noise structure to population state in a single model.
Marcus: It’s an important contribution because it gives us a way to get more honest and mechanistic representations of our biological systems.
Yuki: I think we should definitely keep an eye on this work as we look at how it applies to more complex, real-world ecological datasets moving forward.
Rebecca M. Crossley, Ruth E. Baker
Mathematical Institute, University of Oxford
q-bio.QM, q-bio.PE
Submitted: 2026-06-11
Updated: 2026-10-02
Comments: 32 pages (including two pages SI), 8 figures (including two in SI)
Code: https://github.com/beckycrossley/Learning-noise-growth-BINNs
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: A likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks introduces an extension to existing Biologically-Informed Neural
Key concepts
- Heteroscedastic Noise
- This refers to noise where the variance (or magnitude) is not constant but changes depending on the population density. Instead of assuming all errors are equally likely, this model allows the system to learn that noise becomes larger or smaller as the population grows.
- Noise Scaling Law
- The framework assumes a specific mathematical relationship, a power law ($\sigma(u) = \sigma0|u|\alpha$), to describe how the noise magnitude changes with population density ($u$). The network learns the parameters ($\alpha$ and $\sigma0$) of this law from the data.
- Total Loss Function
- The training process uses a combined loss function consisting of three parts: data likelihood (to fit observations), ODE residual (to ensure dynamics follow the growth equation), and biological constraints (like non-negativity). These are weighted equally to train all aspects simultaneously.
Terminology
Summary
A likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks introduces an extension to existing Biologically-Informed Neural Networks (BINNs) that allows for the direct discovery of a learnable noise model from data. This method is significant because it moves beyond implicitly assuming homoscedastic Gaussian noise, enabling the framework to account for potentially meaningful structure in biological variability by linking noise magnitude directly to the population density.
The gist
The proposed NLL–BINN framework embeds a parametric noise model directly within a likelihood-based formulation of a BINN, allowing both governing dynamics and observation noise structure to be inferred simultaneously from data.
Framework Extension and Formulation
The core extension involves modifying the standard BINN loss function to incorporate heteroscedastic noise. Instead of assuming additive Gaussian noise with constant variance, the framework models observations as:
- Data model: The observed data is modeled using the relationship:
u o i = u(t i) + σ(u(t i))ϵ, where ϵ ∼ N (0, 1).
-
Noise structure: A power-law form is assumed for the noise magnitude to capture density dependence: σ(u) = σ0uα.
-
Learnable parameters: The noise parameters α and σ0 are treated as learnable quantities, enabling the discovery of the noise scaling law directly from observations.
Total Loss Function Components
The total loss function (Equation 5) is a combination of three terms, with equal weighting (wdata = wODE = wbio = 1):
(i) Data Loss (Ldata)
This term is the negative log-likelihood (NLL) loss for the data, based on Equation (3), which accounts for the density-dependent variance:
Ldata = −(1/N) Σ i log [p(u o i uθ(t i))] = 1/N Σ i [u o i − uθ(t i)]2 / [2σ(u o i)2] + 1/2 log[2πσ(u o i)2].
(ii) ODE Loss (LODE)
To ensure consistency with the governing equation, the ODE residual is enforced by randomly sampling NODE time points t ODE ∈ [0, tend] and penalizing the discrepancy:
LODE = (1/NODE) Σ i d/dt uθ(t ODE i) − uθ(t ODE i) gϕ(uθ(t ODE i)).
(iii) Biological Loss (Lbio)
This term enforces soft biological constraints, such as the non-negativity of population densities:
Lbio = (1/Nbio) Σ i max [0, −uθ(t bio i)]2.
Model Evaluation and Results
The framework was tested on synthetic datasets generated from three classical population growth models: the logistic, Gompertz, and Richards’ models. The results demonstrated that the NLL–BINN successfully recovers both the underlying dynamics and the corresponding crowding function g(u) (Figure 1).
-
Accurate Dynamics Recovery: The framework
successfully reconstructs the smooth underlying solution u(t) from noisy data,
with ensemble means closely matching ground truth. -
Noise Identification: The model can correctly identify different forms of heteroscedastic noise; for instance, in the logistic example, it recovers approximately constant variance (α ≈ 0), while in the Richards’ model (multiplicative noise), it recovers linear variance scaling (α ≈ 1).
-
Uncertainty Calibration: The learned uncertainty is well-calibrated across all noise regimes, with observed coverage closely matching theoretical values for multiple coverage levels, achieved
without explicitly enforcing coverage constraints during training.
-
Mechanistic Error Comparison: Compared to a standard BINN trained with RMSE loss, the NLL–BINN achieves a
lower mechanistic error,
indicating that incorporating a learnable noise model improves both the accuracy and robustness of the inferred dynamics.
Biological Application
The framework was applied to real experimental data from coral reef re-growth observations, where variability cannot be estimated from replicated measurements. The method successfully infers both the underlying growth laws and a density-dependent noise model σ(u), showing that the learned noise increases with population density, consistent with heteroscedastic variability observed in these systems. This demonstrates the framework's applicability to real-world data where both dynamics and noise structure are unknown.
Discussion and Implications
The key conceptual contribution is treating noise as a mechanistic quantity rather than a nuisance.
By learning a functional relationship between noise magnitude and population density, the approach provides insight into how uncertainty arises from the underlying system or measurement process.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by incorporating the findings of this NLL-BINN framework:
-
Replacement of fixed homoscedastic noise assumptions with a learnable, density-dependent noise model.
-
Joint learning of system dynamics (ODE) and observation noise structure from noisy data using a likelihood-based formulation.
-
Accurate and interpretable quantification of heteroscedastic uncertainty in mechanistic models, moving beyond simple Gaussian error bounds.
The improved AI system can achieve the following specific capabilities:
-
Predicting population or system states with calibrated uncertainty intervals that correctly reflect density-dependent measurement noise (e.g., additive, multiplicative, or intermediate noise regimes).
-
Automatically inferring the underlying governing growth laws (e.g., logistic, Gompertz, Richards) and their associated parameters directly from sparse or noisy time-series data without requiring prior functional form specification.
-
Identifying and quantifying the specific scaling law of observational variability as a function of the system's state (population density), providing mechanistic insight into why uncertainty is high in certain biological regimes (e.g., low densities or near carrying capacity).
-
Improving mechanistic model recovery by training on observation noise structure rather than just minimizing data misfit, leading to lower
mechanistic error
and more robust inferences about the true underlying dynamics. -
Applying mechanistic learning to real-world experimental data (like coral reef regrowth) where replicated measurements are unavailable, allowing for the simultaneous discovery of both growth laws and noise structures from a single trajectory.
Abstract
In recent years, neural ordinary differential equation frameworks such as Biologically-Informed Neural Networks (BINNs) have shown promise for learning mechanistic laws from sparse data. However, most existing approaches implicitly assume homoscedastic Gaussian noise, and therefore do not account for potentially meaningful structure in biological variability. Here, we present an extension to the existing BINNs framework that includes a learnable noise model, allowing discovery of the noise model directly from data. Using population growth as an example, we demonstrate that the framework accurately recovers the underlying noise structure and improves predictions of the underlying growth laws compared to existing approaches. As such, this work establishes a general likelihood-based framework for jointly learning dynamics and heteroscedastic noise within mechanistic neural network approaches.
Sources
- Functional and parametric identifiability for universal differential equations applied to chemical reaction networks
- Universal Differential Equations for Scientific Machine Learning
Related papers
- Automated Lesion Segmentation of Stroke MRI Using nnU-Net: A Comprehensive External Validation Across Acute and Chronic Lesions
- Resolving satellite-in situ mismatches in Net Primary Production using high-frequency in situ bio-optical observations in the subpolar Northwest Atlantic
- easyplater: The easy way to generate microplate designs deconvolved from multivariate clinical data
- Essential Workers at Risk: An Agent-Based Model (SAFE-ABM) with Bayesian Uncertainty Quantification
- OmniBioTwin: A System-of-Twinned-Systems Framework for Health Digital Twins
- AbFlow: End-to-end Paratope-Centric Antibody Design by Interaction Enhanced Flow Matching