GAME: Genetic Algorithms with Marginalised Ensembles for model-independent reconstruction of cosmological quantities

arXiv:2602.12870 · astro-ph.CO · Submitted 2026-02-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: Next we'll be talking about the paper "GAME: Genetic Algorithms with Marginalised Ensembles for model-independent reconstruction of cosmological quantities".

Jocelyn: The paper was written by M. Peronaci, M. Martinelli and S. Nesseris from University of Rome "Tor Vergata" and Institute National for Nuclear Research (INFN) and University of Rome "Sapienza" and National Institute of Astrophysics - Astronomical Observatory of Rome and Institute of Theoretical Physics (IFT) University of Madrid - CSIC.

Vera: Stay tuned as we take you through the paper and discuss its implications.

Jocelyn: We also have Subrahmanyan with us today — guest researcher.

Vera: Alright, let's get started.

Summary of Findings: Vera: Let’s look at what the paper says it can do, based on its summary, "GAME: Genetic Algorithms with Marginalised Ensembles for model-independent reconstruction of cosmological quantities." The researchers are showing that their method is incredibly effective at reconstructing complex analytical functions directly from data points.

Jocelyn: But they aren't just saying it fits the data; they mention that this technique allows us to look beyond simply finding the minimum chi two value, which is a huge step for researchers who are trying to understand the real behavior of things.

Subrahmanyan: The authors are demonstrating that this ensemble approach significantly reduces the uncertainty caused by running an AI algorithm, making it a very stable tool for modeling complex cosmic evolution.

Vera: And when they talk about applying this method to Cosmic Chronometers data, they're showing how we can finally get a reliable picture of the expansion rate H(z) without being overly reliant on one specific parameter set.

Jocelyn: It sounds like the initial results are already promising, confirming that the method is able to reconstruct H(z) in a way that matches what we expect from our current understanding of dark energy.

Subrahmanyan: The paper's summary suggests that we have a powerful, new tool to handle the stochastic nature of data analysis while maintaining genuine scientific rigor.

Vera: This really shows how much more robust this whole process is compared to just running a standard genetic algorithm and trusting the only best output you get from one single run.

Jocelyn: It's clear that moving past the limitations of a single run opens up possibilities for deeper analysis, so let's see how they actually build this into the next segment.

Improvements in Methodology: Jocelyn: Now, let’s look deeper into the mechanics of "GAME: Genetic Algorithms with Marginalised Ensembles for model-independent reconstruction of cosmological quantities." The researchers have put together several very sophisticated ways to refine this model-averaging approach.

Vera: It's not just a simple averaging; it’s about choosing a smart weighting scheme, which involves a combination of the goodness-of-fit and a smoothness penalty, which they call R j.

Subrahmanyan: That roughness term, R j, is what mathematically quantifies how well we are keeping the first derivative consistent across all those different configurations generated by the Genetic Algorithm.

Jocelyn: And to decide which configurations are worth including in that average, they use something called an L-curve method to determine a regularisation parameter, lambda. It’s essentially a way for us to choose the right level of caution.

Vera: It's not just about minimizing the error; it’s about finding the optimal balance between fitting the data perfectly and ensuring that the resulting function is stable enough to yield reliable derivatives.

Subrahmanyan: The L-curve provides a rigorous, visual way to select that optimal trade-off point, where we transition from simply fitting random noise to actually capturing physical signal.

Jocelyn: They also provide a practical way to rigorously estimate errors on this averaged function by combining a path-integral approach with an ensemble variance, which is something standard GA simply cannot do.

Vera: This means the uncertainty isn't just one static number; it accounts for how much the different GA models are diverging from each other, which is crucial when we need stable derivatives.

Subrahmanyan: By quantifying that ensemble spread, they are essentially giving us a measure of the inherent uncertainty in running an AI simulation itself.

Jocelyn: That level of detail is reassuring, Subrahmanyin; it makes the whole system feel much more trustworthy before moving on to see what these tools can do with real data.

Results and Impact: Vera: We've seen the technical improvements in "GAME: Genetic Algorithms with Marginalised Ensembles for model-independent reconstruction of cosmological quantities," so let's look at what this means for real data using Cosmic Chronometers.

Jocelyn: The key finding is that by using GAME, they can derive the dark energy equation of state w(z) while remaining consistent with the expected model at low redshifts.

Subrahmanyan: This is significant because the paper shows that even with current data, this approach provides a much more reliable picture than what we got before using traditional methods.

Vera: And they are looking ahead, testing the method against mock data from Stage IV surveys to show it works for future precision cosmology where our current constraints are limited by measurement uncertainty.

Jocelyn: The forecast simulations confirm that as survey precision increases, GAME is ready to provide much tighter constraints on w(z) than we can achieve today with our current telescopes.

Subrahmanyan: It’s a clear demonstration that the statistical power of this method allows us to probe the expansion history of the universe in a much more robust way than previously possible.

Vera: This is not just about fitting data; it' about ensuring that we are generating results that are scientifically sound, regardless of how many different ways those results could have been achieved.

Jocelyn: The potential for future observations is really what makes this exciting; the predictive power of this method opens up a whole new era.

Conclusion: Vera: We’ve seen so much ground today discussing "GAME: Genetic Algorithms with Marginalised Ensembles for model-independent reconstruction of cosmological quantities," and I think the implications for our field are truly profound.

Jocelyn: The way this method handles both the statistical measurement uncertainty and the configuration-driven spread means that we're ready to use these high-precision Stage IV surveys with confidence.

Subrahmanyan: I think the mathematical rigor behind weighting those results with a roughness penalty is what ensures that our derived quantities, like w(z), are not just fitting noise but actually capturing a smooth, physical reality.

Vera: That stability in the data is exactly what we're looking for when trying to constrain models; you can't have reliable derivatives if the function itself is oscillating wildly at high redshifts.

Jocelyn: And I agree with Vera; seeing that consistency across multiple simulations suggests that this tool is ready for the precision of Stage IV surveys, which will be observing us with unprecedented detail.

Subrahmanyan: It’s a testament to the model-independent nature of this approach, allowing us to let the data guide our understanding rather than forcing it into some pre-selected theoretical box.

Vera: I'm excited to see how this translates into actual discovery, knowing that we have these robust methods to check for any subtle deviations from.

Jocelyn: It’s a huge relief that the uncertainty analysis is so well-defined; it won't just be a single number, but a total confidence band incorporating both measurement noise and the stochastic spread.

Subrahmanyan: This allows us to confidently report that our findings are robust against overfitting, which is absolutely critical when we are trying to establish new physics.

Vera: I hope this opens the door for even more rigorous model comparisons in future work, allowing us to see what really is out there beyond our current expectations.

Jocelyn: I think so too; we're really looking forward to applying this technique as we start processing the data from these next generation telescopes.

Subrahmanyan: It provides a solid framework for finding answers, even if those answers are surprising us in new ways that challenge our current understanding.

Vera: We'll wrap up our talk on this groundbreaking work now, and I know the community is going to be very interested in what’s next for cosmology using "GAME: Genetic Algorithms with Marginalised Ensembles for model-independent reconstruction of cosmological quantities."

University of Rome "Tor Vergata" · Institute National for Nuclear Research (INFN) · University of Rome "Sapienza" · National Institute of Astrophysics - Astronomical Observatory of Rome · Institute of Theoretical Physics (IFT) University of Madrid - CSIC

astro-ph.CO

Submitted: 2026-02-13

Updated: 2026-09-03

Comments: 29 pages, 13 figures. Accepted for publication in JCAP

Journal ref: JCAP 08 (2026) 063

DOI: 10.1088/1475-7516/2026/08/063

Code: https://github.com/snesseris/Genetic-Algorithmshttps:

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: The paper introduces GAME (Genetic Algorithms with Marginalised Ensembles), a sophisticated computational framework designed for the model-independent reconstruction of fundamental cosmological

Key concepts

Model-independent reconstruction
This technique reconstructs complex physical functions, such as the universe's expansion rate $H(z)$, directly from observational data. It allows researchers to let the data guide understanding rather than forcing results into a pre-selected theoretical model.
Marginalised Ensembles
This statistical approach uses multiple runs of an AI algorithm to calculate uncertainty. Instead of a single error number, it quantifies the spread between different model outcomes, providing a robust measure of inherent uncertainty in the data analysis process.
Smoothness Penalty ($R_j$)
This is a mathematical term used to ensure physical stability in the reconstructed function. It penalizes rapid changes or inconsistencies in the function's derivatives, ensuring that derived quantities are reliable and physically plausible.

Terminology

Summary

The paper introduces GAME (Genetic Algorithms with Marginalised Ensembles), a sophisticated computational framework designed for the model-independent reconstruction of fundamental cosmological quantities, such as the Hubble rate H(z) and the dark energy equation of state w(z). This methodology is critical because it allows researchers to extract physical parameters directly from observational data without having to assume a specific underlying cosmological model (like CDM), thereby providing stringent tests for standard cosmology using diverse datasets, particularly Cosmic Chronometers (CC).

Hyperparameter Configuration and Optimization

The reconstruction process relies on Genetic Algorithms (GA) applied over an ensemble of configurations. The hyperparameters used are meticulously controlled for both the validation test function f(x) and the CC reconstruction of H(z). For instance, the GA setup involves:

  • Maximum generations (N gen) and population size (N pop), which were fixed at 3000 and 100, respectively.

  • The genetic operators rates (selection rate, crossover rate, mutation rate) were set to 0.30, 0.85, and 0.85.

  • The methodology considers multiple grammar sets for the reconstruction, including combinations like [poly], [cpl], and [poly, cpl] for f(x), demonstrating a rigorous search over functional forms.

Impact of Data Quality on Reconstruction Robustness

The primary analysis utilized the full available dataset of CC up to z about 2, acknowledging that the CC dataset is heterogeneous, combining measurements from various surveys with different sensitivities and systematics. To verify the impact of lower-precision data, an alternative run was performed using a quality-cut subset. This subsample excluded measurements where H > H mean + 2 sigma H, thereby removing points susceptible to large systematic effects or low signal-to-noise ratios.

**Reconstruction of H(z) and its Derivative H'(z) **

When comparing the full dataset analysis to the high-quality subset, the reconstruction of H(z) was found to be fully consistent with the full-dataset analysis, confirming that the GAME approach is robust to the exclusion of low-quality measurements. However, the improvement is most notable in the derivative H'(z). In the main analysis, H'(z) exhibited fluctuations at high redshifts (z > 1.2) due to sparser and noisier data. By removing outliers, the resulting H'(z) becomes significantly smoother and tracks CDM prediction more closely, suggesting that the slight oscillatory features in the full analysis were likely artifacts driven by the scatter of low-precision data points rather than physical deviations from the standard model. Furthermore, this exclusion decreases the pathintegral uncertainty estimation delta f PI.

**Derivation and Constraints on Dark Energy w(z) **

The smoothness of H'(z) directly impacts the derivation of w(z). The use of the quality-cut subset leads to a slightly mitigated divergence of confidence intervals for w(z) at high redshifts (z > 1.5). Crucially, the constraints at low redshift (z < 1) remain tight and centered on w = -1. When adopting the DESI prior for m,0, the resulting constraint was found to be stronger than that derived using the full sample: w(0) = -0.988 plus or minus 0.101. This test demonstrates that while all available data is statistically complete, the future of model-independent reconstructions lies in the acquisition of high-precision CC measurements rather than simply increasing the data sample.

Improvements for AI systems

The scientific paper details complex analyses involving model-independent reconstruction of cosmological parameters (H(z), w(z)) using heterogeneous and systematically biased datasets (Cosmic Chronometers, CC). The core challenges are: 1) Fusing low signal-to-noise, high-uncertainty data streams; 2) Performing robust inference while respecting physical priors (CDM); and 3) Handling complex parameter space exploration (Genetic Algorithm setup).

I propose three major improvements targeting Data Fusion/Noise Robustness, Uncertainty Quantification, and Model Selection/Prior Enforcement.


Current Limitation Addressed: The standard chi squared optimization (as used in the GAME pipeline) struggles with highly heterogeneous data streams where different surveys have varying systematic errors, intrinsic noise profiles, and differing levels of constraining power. Simple averaging or weighting fails to capture the full covariance structure.

Proposed AI Improvement: Implement a Hierarchical Bayesian Deep Learning (HBDL) architecture structured as a combination of Variational Autoencoders (VAEs) and Recurrent Neural Networks (RNNs), specifically utilizing a Transformer layer for long-range dependency modeling.

  1. Input Layer: Each CC measurement i at redshift z i is treated as an input vector x i = H obs, sigma stat, sys.

  2. Encoding/Feature Extraction (VAE): A VAE processes the raw data to learn a low-dimensional, noise-robust latent representation z i for each measurement. The VAE explicitly models the uncertainty (sigma total) as part of the latent space, rather than just predicting a mean value.

  3. Fusion/Inference (Transformer/RNN): The sequence of latent representations z 1, z 2,..., z N is fed into a Transformer block with masked self-attention. This allows the model to dynamically weight the influence of measurements based on their relative reliability and temporal proximity, effectively learning a data-dependent covariance matrix data(z).

  4. Output: The system outputs the reconstructed H(z) and its full posterior probability distribution P(H(z) Data), which inherently incorporates the systematic uncertainty structure (sys) from all sources.

Improved AI Capability:

The system can generate a **Dynamically Weighted, Systematically Corrected Reconstruction of H(z) **. It will automatically identify and downweight measurements that exhibit statistically anomalous behavior (outliers) relative to the global trend without requiring manual thresholds (e.g., H > H mean + 2 sigma H), thereby surpassing the limitations of simple quality-cutting methods while providing a full, continuous posterior distribution for H(z) and H'(z).

  1. Architecture: The PINN is designed not just to fit the data points H(z), z, but to simultaneously minimize a loss function that includes three components:

L total = L data + lambda 1 L priors + lambda 2 L physics

  1. Data Loss (L data): Standard mean squared error between the network output and the reconstructed H(z) values.

  2. Prior Loss (L priors): Enforces known cosmological constraints (e.g., Planck m,0, or a specific DESI prior) by penalizing deviations from these established ranges in the latent space.

  3. Physics Loss (L physics): This is the critical addition. It embeds the underlying differential equations governing cosmology (e.g., Friedmann equations relating H(z) to m,) directly into the loss function via automatic differentiation (d L over d w). This forces the network's output trajectory H NN(z) to satisfy known physical laws throughout the entire redshift range, even where data is sparse.

  4. Graph Construction: Define the hyperparameters (e.g., N gen, Selection Rate, Grammar Choice) as nodes in a graph G. The edges represent the dependency or interaction between these hyperparameters (e.g., how N pop influences the required N gen).

  5. Feature Embedding: Each node is initially embedded with data from previous GA runs (results, convergence speed, final fitness).

  6. Inference/Optimization: The GNN processes this graph structure to predict the optimal configuration of hyperparameters for a new target quantity (e.g., H(z) vs. f(x)). Instead of running the full GA suite blindly, the GNN guides the search by predicting which combinations of grammar and rates are most likely to achieve convergence within a specified computational budget.

Sources

Related papers