Simulation-based inference from the Lyman-alpha forest 1D power spectrum with CAMELS

arXiv:2603.13011 · astro-ph.CO · Submitted 2026-03-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: Today's paper: "Simulation-based inference from the Lyman-alpha forest 1D power spectrum with CAMELS".

Jocelyn: Simulation-based inference from the Lyman-alpha forest 1D power spectrum with CAMELS performs full simulation-based inference on the Lyman-alpha forest 1D power spectrum,

Vera: First, who's behind it and why it matters.

Paper summary: Vera: Well, Jocelyn, we're starting with the paper "Simulation-based inference from the Lyman-alpha forest 1D power spectrum with CAMELS <ref:2603.13011#pg0,Simulation-based inference from the Lyman-alpha forest 1D power spectrum with>." It's pretty interesting because they perform full simulation-based inference on the Lyman-alpha forest 1D power spectrum using CAMELS <ref:2603.13011#pg0,full simulation-based inference on the Lyman>. The main idea is that they train a normalizing flow to estimate two cosmological parameters and four astrophysical parameters. This approach is significant because it tackles the problem of getting convergence from different galaxy formation models.

Jocelyn: That sounds like a really sophisticated way to move beyond traditional likelihood methods, Vera. So, what's the core claim here regarding why this matters for cosmology and physics? The summary suggests it offers a promising new path forward using artificial intelligence in constraining these parameters.

Subrahmanyan: I agree that the focus on overcoming issues with different galaxy formation models is crucial; if those models don't converge, our cosmological constraints become shaky. This paper claims this AI approach provides a way to constrain cosmology and fundamental physics using the Lyman-alpha forest data, which is a powerful probe of the cosmic web.

Vera: Exactly, Subrahmanyan. The paper states they are using CAMELS simulations run with IllustrisTNG and SIMBA galaxy formation models for this inference. They are specifically looking at two cosmological parameters: m and sigma eight and four astrophysical parameters that describe supernova and AGN feedback.

Jocelyn: So, the method relies on training a neural network to map the Lyman-alpha forest power spectrum directly to those parameter sets, instead of relying solely on likelihoods derived from simulations. It's about learning the mapping between the observable data and our physical inputs.

Subrahmanyan: That's a key distinction; they are using Simulation-Based Inference, which learns this mapping from the simulations themselves rather than just fitting them with traditional likelihood functions <ref:2603.13011#pg1>. This has potential because it can handle complex degeneracies that might confuse simpler analytical approximations <ref:2603.13011#pg2>.

Vera: And the training setup is quite rigorous, as they trained the network on nine hundred fifty simulations—four hundred seventy-five from IllustrisTNG and four hundred seventy-five from SIMBA—and then tested its performance on another set of fifty simulations. That cross-validation helps establish how robust this AI method is.

Jocelyn: It sounds like the paper is setting up a comprehensive comparison between two different galaxy formation frameworks, IllustrisTNG and SIMBA, to see if the AI can truly bridge the gap between them. This is where I'm curious about the practical implications for future surveys we might run.

Subrahmanyan: Precisely, Jocelyn; when they test on one model and train on another, they see performance drop significantly, resulting in about a ten percent positive bias on sigma eight. However, their finding that multi-domain training based on combining simulations from both models recovers unbiased constraints is a strong statement about the power of this AI framework.

Vera: That recovery of unbiased constraints by combining both models is what makes this work so important for constraining cosmology and fundamental physics using the Lyman-alpha forest data. It tackles that lack of convergence we talked about earlier.

Paper summary: Jocelyn: So, if we look at the scale of the results they achieved, they report that for IllustrisTNG, inference predicts unbiased values within ten percent or twenty percent deviations in a significant portion of cases for m and sigma eight. That level of precision is quite compelling when you consider the inherent noise in observational data.

Subrahmanyan: Indeed, the precision achieved is about eight percent for m and about six percent for sigma eight when training on IllustrisTNG, which suggests a solid footing for constraining those cosmological parameters using this method <ref:2603.13011#pg0>.

Vera: It’s exciting to see how these simulations can be leveraged with AI to extract so much information from the Lyman-alpha forest P1D(k) measurements. But we do have to keep an eye on the limitations they mention, which is that the CAMELS suite explores only two cosmological parameters, m and sigma eight.

Jocelyn: And I also noted their comment about the simulations themselves having fairly low resolution and small volume. That's a practical constraint we need to keep in mind when thinking about how far these constraints can actually push our understanding of the universe.

Subrahmanyan: They explicitly limit their baseline analysis to conservative scale cuts, specifically k max values of one point five h Mpc-one two point zero h Mpc-one and three point zero h Mpc-one because they needed stable and converged results given the resolution limitations of the training simulations <ref:2603.13011#pg2>.

Vera: So, to recap, this paper shows that using AI to infer cosmological parameters from the Lyman-alpha forest power spectrum via simulation-based inference can provide unbiased constraints by intelligently combining outputs from different galaxy formation models. The main implication is that we might finally be able to use this specific observable with greater confidence for testing cosmological models and understanding the physics driving structure formation.

Jocelyn: It really does feel like a step forward in how we can process complex astrophysical simulations into something usable for parameter estimation, Vera. It opens up new avenues for combining observational data with these complex simulation results.

Subrahmanyan: I think the real impact here is demonstrating that this AI methodology doesn't just offer better fits but provides a way to navigate the complexity introduced by varying galaxy formation physics without losing cosmological accuracy <ref:2603.13011#pg0>.

Vera: So, while we have these exciting initial results, what does this mean for future work? The paper mentions exploring other redshifts and potentially testing more complex parameter spaces in the future.

Jocelyn: I'm looking forward to seeing how this framework evolves; it feels like the starting point for a whole new class of analyses using AI on cosmological simulations.

Subrahmanyan: The next steps will involve pushing these constraints further, perhaps by incorporating more astrophysical parameters or exploring a wider range of cosmological inputs than just m and sigma eight though that requires more robust simulation suites <ref:2603.13011#pg2>.

Vera: It sounds like a promising direction for observational cosmology, Jocelyn, showing us how we can harness advanced computational tools to tackle the messy reality of galaxy formation modeling.

Conclusion: Vera: So, to recap, this paper introduces a new way to use AI to figure out cosmological and astrophysical parameters by looking at the Lyman-alpha forest power spectrum from simulations. Jocelyn, what do you think about that title?

Jocelyn: I think the title is pretty descriptive; it immediately tells us what observable they're using—the Lyman-alpha forest—and how they’re getting their answers, which is through simulation-based inference. It sounds like a solid piece of work for anyone studying structure formation.

Subrahmanyan: From a theoretical standpoint, the focus on simulation-based inference suggests a real push toward making our cosmological models more tractable by linking them directly to complex hydrodynamical simulations. This moves us beyond just fitting simplified analytical models and forces us to confront the full complexity of galaxy formation physics.

Vera: Exactly, Subrahmanyan. And looking at the authors, I see a solid team behind this effort, which speaks to how much computational power is needed for these kinds of large-scale comparisons. It shows that this kind of work isn't just one person's idea but a collaborative push.

Jocelyn: And those authors are clearly tackling some tough problems in the field, like getting convergence from different galaxy formation models, which sounds like a major hurdle they’re trying to clear for us observationalists. It’s exciting to see them address that difficulty head-on.

Subrahmanyan: That difficulty is precisely where the real physics lies; if we can consistently constrain m and sigma eight across different feedback prescriptions, it gives us much firmer ground for testing how dark matter clusters and how early structure formed in the universe.

Vera: And that's what I find so compelling from an observational side; having better constraints on these parameters means we can put tighter limits on the underlying physics driving the evolution of galaxies across cosmic time. It really connects what we see in the sky to what’s happening inside those simulated boxes.

Jocelyn: So, this paper essentially shows us a new toolkit—one that uses AI to sift through massive simulation data and pull out meaningful cosmological constraints without getting bogged down by the fine details of every single astrophysical sub-process. That capability is really what makes this research interesting for our field.

Subrahmanyan: It opens up a path where we can systematically probe the interplay between dark matter gravity, cosmology, and the complex feedback mechanisms that dictate how galaxies actually assemble their stars and gas. We're talking about getting closer to understanding the cosmic web in its entirety, not just a simplified snapshot.

Institute for Fundamental Physics of the Universe (IFPU) · SISSA - International School for Advanced Studies · INAF - Osservatorio Astronomico di Trieste · INFN – National Institute for Nuclear Physics

astro-ph.CO

Submitted: 2026-03-13

Updated: 2026-10-06

Comments: 37 pages, 27 figures. Accepted for publication in Physical Review D

Code: https://github.com/sbird/fake_spectra

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 78/100

The gist: Simulation-based inference from the Lyman-alpha forest 1D power spectrum with CAMELS performs full simulation-based inference on the Lyman-alpha forest 1D power spectrum, training a normalizing flow

Key concepts

Simulation-Based Inference (SBI)
SBI is an AI technique that learns the relationship between simulation inputs and desired outputs without needing a traditional mathematical likelihood function. It trains a neural network to predict parameter values directly from simulated data, offering an alternative way to constrain physics.
Lyman-alpha forest 1D power spectrum (P1D(k))
This is the observable data used in the study. It measures how the density fluctuations of neutral hydrogen gas are distributed across different spatial scales (represented by 'k'). Analyzing this spectrum allows researchers to probe the underlying cosmological and astrophysical properties of the universe.
Normalizing Flow (MAF)
A specific type of neural network architecture used for SBI, Masked Autoregressive Flows. This model is trained to map the observed P1D(k) data into a distribution that represents the posterior probability of the model parameters, allowing for precise estimation of uncertainties.

Terminology

Summary

Simulation-based inference from the Lyman-alpha forest 1D power spectrum with CAMELS performs full simulation-based inference on the Lyman-alpha forest 1D power spectrum, training a normalizing flow to perform neural posterior estimation of two cosmological parameters and four astrophysical parameters. This work establishes a promising way forward to constrain cosmology and fundamental physics with the Lyman-α forest using artificial intelligence by overcoming the lack of convergence in predictions from different galaxy formation models.

The Gist

We perform for the first time full simulation-based inference on the Lyman-α forest 1D power spectrum, considering two cosmological parameters (omegam and σ8) and four astrophysical parameters parametrizing supernova and AGN feedback, using CAMELS cosmological hydrodynamic simulations run with IllustrisTNG and SIMBA galaxy formation models.

Simulation Setup

The study utilizes the CAMELS suite, which comprises a large set of state-of-the-art cosmological N-body and magneto-hydrodynamic simulations. These simulations were obtained using two different codes and baryon physics models: (i) one set run with AREPO and the IllustrisTNG model, and (ii) one set run with the GIZMO code and the SIMBA model. The simulations were run in cosmological boxes of volume V = (25 h−1 Mpc)3 with periodic boundary conditions, following the evolution of 2563 dark matter particles of mass mdm = 6.49 × 107 (omegam − omegab)/0.251 h−1 M⊙ and of 2563 gas resolution elements with an initial mass mgas = 1.27 × 107 h−1 M⊙. The CAMELS LH simulations explore the parameters:

** Cosmological parameters: omegam ∈ [0.1, 0.5], σ8 ∈ [0.6, 1.0].**

** Astrophysical parameters: ASN1, AAGN1, ASN2, AAGN2 sampled following a latin hypercube strategy with priors.**

Inference Framework

The paper employs Simulation-Based Inference (SBI), which is an alternative to traditional likelihood-based analysis that learns the mapping between parameters and observables directly from simulations. Specifically, they train a neural network to approximate the posterior distributions of the model parameters given the Lyα forest P1D(k). They adopt a normalizing flow model, specifically Masked Autoregressive Flows (MAF), which consists of a stack of autoregressive models. The network is trained on 950 simulations (475 from IllustrisTNG, 475 from SIMBA) and tested on 50 simulations (25 from IllustrisTNG, 25 from SIMBA). The input to the neural network is a one-dimensional data vector obtained by stacking the P1D(k) measurements at the 8 different available redshifts.

Results and Discussion

The study evaluates performance by measuring accuracy and precision of predicted posteriors compared to true parameter values.

** Training on the same galaxy formation model (IllustrisTNG or SIMBA) yields unbiased constraints on the cosmological parameters, but leaves the astrophysical ones unconstrained, likely due to volume effects.**

** When training on one model and testing on another (e.g., IllustrisTNG and SIMBA), performance is significantly worse, resulting in a ∼ 10% positive bias on the predicted values for σ8.**

** A multi-domain training based on the combination of simulations from both models recovers unbiased constraints, offering an effective solution to cope with the complex problem of the lack of convergence in predictions from different galaxy formation models.**

The results show that for IllustrisTNG, inference predicts unbiased parameter values within 10% (20%) deviations in ≳ 75% (100%) of the cases for omegam and in ≳ 90% (100%) of the cases for σ8. The precision achieved is ∼ 8% and ∼ 6% on these parameters, respectively. When training on SIMBA, the results are slightly worse than those for IllustrisTNG, potentially reflecting a stronger dependence on astrophysical parameters which reflects into a larger uncertainty.

Limitations and Future Work

The CAMELS LH suite has limitations: it explores only two cosmological parameters (omegam and σ8), and the simulations have fairly low resolution and small volume. The study limits its baseline analysis to conservative scale cuts—kmax = 1.5 h Mpc−1, kmax = 2.0 h Mpc−1, and kmax = 3.0 h Mpc−1—to ensure stable and converged results due to the limited resolution of the training simulations.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed this cutting-edge paper on simulation-based inference (SBI) of the Lyman-alpha forest power spectrum. The work presents a novel framework that leverages normalizing flows (specifically Masked Autoregressive Flows, MAF) to perform neural posterior estimation for cosmological parameters from complex hydrodynamical simulations.

Here are the specific improvements I can suggest for AI systems, categorized by the technical aspect they target:


) Improvements to AI Systems Derived from This Paper:

  1. A new class of Machine Learning (ML) inference algorithms specifically tailored for high-dimensional, non-Gaussian posterior estimation in cosmological contexts.

  2. A robust framework for Domain Adaptation in simulation-based inference, allowing models trained on one physical code (e.g., IllustrisTNG) to be reliably transferred and tested on outputs from a different code (e.g., SIMBA).

  3. A method for mitigating the Cross-Generalization Problem in SBI by explicitly learning shared patterns across diverse astrophysical physics models, ensuring that inferred cosmological constraints are not artifacts of the specific simulation suite used for training.

) What the Improved AI System Can Do (Specific Capabilities):

  1. A new ML inference system capable of taking a set of 1D flux power spectrum measurements from the Lyman-alpha forest as input and directly outputting high-precision, marginalized posterior distributions for cosmological parameters like Matter Density and Growth Rate, even when the underlying physics is complex (i.e., non-linear structure formation).

  2. This system can perform Parameter Space Exploration by rapidly sampling the entire posterior distribution of multiple cosmological parameters simultaneously, offering a way to explore high-dimensional parameter spaces without requiring expensive traditional Markov Chain Monte Carlo (MCMC) methods.

  3. The improved Domain Adaptation system allows an AI model to be trained on a diverse set of galaxy formation simulations (IllustrisTNG and SIMBA). The resulting AI can then be reliably used for inference on unseen simulation outputs, effectively reducing the dependency on perfectly converged predictions from any single simulation code.

  4. The system can provide a diagnostic tool that assesses the reliability of inferred cosmological constraints by calculating posterior coverage tests (like TARP), ensuring that the reported uncertainties are statistically sound and not overly optimistic (i.e., preventing overconfidence).

  5. By utilizing Optical Depth instead of raw flux, the system can be adapted to infer cosmology even when instrumental noise or continuum subtraction effects are present, leading to a more robust measurement of cosmological parameters than direct flux estimation alone.

Abstract

We perform for the first time full simulation-based inference on the Lyman- α forest 1D power spectrum. In particular, we consider the prediction of the Lyman- α forest P 1D(k) at 2.0<z<3.5 from the CAMELS cosmological hydrodynamic simulations run with the IllustrisTNG and SIMBA galaxy formation models. We train a normalizing flow to perform neural posterior estimation of two cosmological parameters (Ω m and σ 8) and four astrophysical parameters parametrizing supernova and AGN feedback. When training and testing the neural network on the same baryon physics model, the posterior distributions of the cosmological parameters are found to be in excellent agreement with the true parameters values (within 10% deviations in 75% and 90% of the cases for Ω m and σ 8, and a precision better than 10% in both), while the astrophysical parameters converge to the prior mean due to the limited probed volume. When training on one model and testing on the other (e.g., training on IllustrisTNG and testing on SIMBA, or viceversa), the performance is significantly worse, both in accuracy and in precision, resulting in a about 10% positive bias on the predicted values for σ 8. We show that a multi-domain training based on the combination of simulations from both models recovers unbiased constraints, offering an effective solution to cope with the complex problem of the lack of convergence in the predictions from different galaxy formation models. This study represents a promising way forward to constrain cosmology and fundamental physics with the Lyman- α forest with artificial intelligence.

Sources

Related papers