Euclid preparation. CosmoPostProcess: A simulation calibrated framework for weak lensing selection bias in richness-selected galaxy clusters

arXiv:2605.02723 · astro-ph.CO, astro-ph.IM · Submitted 2026-05-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: Today's paper: "Euclid preparation. CosmoPostProcess".

Jocelyn: The paper details advanced methodological frameworks—including sophisticated statistical calibration and machine learning emulators—designed to accurately quantify and correct for selection bias in richness-selected galaxy clusters.

Vera: First, who's behind it and why it matters.

Title and authors: Vera: So we've covered the calibration and the emulator training; now let's go through what the paper actually summarizes about CosmoPostProcess and what it means in practical terms for our work.

Jocelyn: I think the summary focuses on how CosmoPostProcess acts as a simulation-based forward model algorithm that is specifically calibrated to reproduce optical cluster observables exactly as they are measured by the Euclid space telescope.

Subrahmanyan: Essentially, the main deliverable of CosmoPostProcess is this correction for stacked surface-density profiles, which are binned in richness and redshift, that accounts for selection-related systematic effects.

Vera: So it’s summarizing that this correction takes into account how the stacked weak lensing signal gets modified when we use richness-selected samples identified in the photometric Euclid survey compared to an unbiased reference sample.

Jocelyn: That means they are summarizing the core function: correcting the measured stacked profiles to reflect what they would look like if we had a truly unbiased sample.

Subrahmanyan: They emphasize that this correction is particularly focused on the Euclid richness definition currently foreseen for cosmological analysis, which doesn't involve any color selection in this context.

Vera: That’s a vital detail because it confirms that the calibration is tailored precisely to the specific richness definition we are aiming to use for our cosmology.

Jocelyn: So, in simple terms, they’re telling us that by applying this correction, we get a corrected measurement of cluster density profiles that accurately reflects the underlying matter distribution.

Subrahmanyan: They are showing how this process helps in overcoming the difficulty of using richness as a proxy for mass by modeling and quantifying those selection effects through simulation calibration.

Vera: It sounds like the summary is emphasizing that they are providing a concrete, calibrated tool to bridge the gap between what we observe optically and what we need for cosmology.

Jocelyn: Exactly; so it boils down to this: if you use their method, you get measurements of cluster mass that are better grounded in reality because they have had the selection biases systematically accounted for.

Subrahmanyan: This is how they aim to ensure that the cosmological constraints we derive from Euclid data are not tainted by observational artifacts related to how we select our clusters.

Vera: It really highlights the importance of this work in ensuring the quality of the input data is top-notch before we even start looking for cosmological signals.

Jocelyn: So, moving on, what about the next step in their approach to improving these existing methods? What kind of improvements are they suggesting for future work?

Subrahmanyan: The paper suggests moving toward more advanced AI techniques to enhance the membership emulator and statistical modeling by transitioning from standard fully connected networks to Graph Neural Networks and using Bayesian Deep Learning.

Vera: That means they are looking beyond the current setup, trying to model spatial dependencies in the data structure more deeply with GNNs, which is a natural next step for complex astrophysical data like galaxy clusters.

Jocelyn: And they also mention using Normalizing Flows to model the full conditional probability mass function directly instead of relying on fixed parametric forms for likelihood estimation.

Subrahmanyan: By using Normalizing Flows, they are aiming to make the statistical modeling less dependent on strong assumptions about the underlying physical process, allowing it to handle complex distributions more flexibly.

Vera: So they’re essentially proposing a path toward a more flexible and powerful inference engine that can adapt better to the messy reality of astrophysical observations.

Jocelyn: That sounds like they are building something that is designed not just to fit the data, but to understand the underlying physics through a more nuanced statistical lens.

Subrahmanyan: The ultimate goal is achieving a system where we can provide full predictive uncertainty envelopes for every single cluster detection based on its observational properties.

Vera: That level of detail in uncertainty quantification is exactly what makes the final output trustworthy enough for serious cosmological inference.

The paper's summary: Jocelyn: Okay, we’ve established the core findings, so now let's talk about the specific methodological upgrades the authors propose to improve their framework. What kind of enhancements are they suggesting for future development?

Vera: The paper suggests moving away from standard fully connected networks for P mem toward Graph Neural Networks, specifically a Graph Attention Networks, because that architecture can inherently model the spatial relationships in the data structure better.

Subrahmanyan: A GAT would be advantageous because it lets us assign different weights to different neighbors based on their physical relevance, which is a significant improvement over treating features as independent vectors.

Jocelyn: And what about the data representation itself? They suggest incorporating specialized encoding for geometric constraints, especially since we're dealing with projected separation r/r cl and redshift z.

Vera: They specifically suggest using methods like Spherical Harmonics or Wavelet Transforms on angular coordinates if available, to capture large-scale cosmic structure information more efficiently than just simple linear input features.

Jocelyn: That sounds like a way to inject the geometry of the survey volume directly into the model, making it better at handling varying field geometries.

Subrahmanyan: From a theoretical perspective, they also propose using Bayesian Deep Learning to map catalogue features directly to the posterior distribution parameters instead of just point estimates.

Vera: Instead of just giving us one number for alpha or M one BDL would allow us to see the full range of possibilities for those parameters from the data.

Jocelyn: And they also propose using Normalizing Flows to model the full conditional probability mass function directly, which lets the model learn the shape of that probability distribution without assuming a fixed functional form.

Subrahmanyan: This approach allows us to handle complex and multimodal distributions governing cluster richness more flexibly than a standard negative binomial likelihood.

Vera: So, they are pushing for a system that is not just accurate but also flexible enough to capture the full complexity of the underlying physics without being locked into a single mathematical assumption.

Jocelyn: It sounds like they're aiming to create an inference engine that can handle the inherent noise and diversity in cluster populations more gracefully.

Subrahmanyan: The hope is that this evolution allows us to move toward providing full predictive uncertainty envelopes for every single cluster detection based on its observational properties.

Vera: That level of detail in uncertainty quantification is exactly what makes the final output trustworthy enough for serious cosmological interpretation.

The paper's improvements: Jocelyn: So, we've walked through the paper 'Euclid preparation. CosmoPostProcess: A simulation calibrated framework for weak lensing selection bias in richness-selected galaxy clusters', covering everything from the initial calibration to the future architectural suggestions. Let’s summarize the main implications before we wrap up.

Vera: The main implication is that this work gives us a calibrated framework to handle selection bias in richness-selected galaxy clusters, which is essential for preparing our Euclid data for cosmological analysis.

Subrahmanyan: This means we can start extracting robust constraints on dark energy and the geometry of the universe using weak lensing by having properly accounted for the observational biases introduced by cluster richness selection.

Jocelyn: So, in a nutshell, they’ve delivered a tool that ensures our cosmological results from Euclid are not based on flawed input data because it has systematically corrected for those biases.

Vera: It really puts us in a position to trust the inputs and start building the next generation of cosmological models with confidence.

Jocelyn: I think we’ve got a clear picture now of how this framework helps bridge the gap between observation and theory, so we can finally start using this corrected data for Euclid science.

Subrahmanyan: Ultimately, by providing these calibrated tools, they are enabling us to connect the observational measurements to the big picture of structure formation in a way that is far more reliable than previous methods allowed.

Vera: It’s been fascinating following this paper as we've seen how complex modeling can be applied effectively to tackle real observational challenges.

Jocelyn: We've definitely got some exciting ideas for where the next research efforts should focus, and it’s clear that the path toward using these sophisticated techniques is very promising.

Subrahmanyan: I just think this work sets a strong foundation for how we connect galaxy cluster counts to cosmological evolution in a way that is far more reliable than before.

Conclusion: Jocelyn: So, to wrap up this discussion on "Euclid preparation: CosmoPostProcess," we’ve seen how they used simulation calibration to correct selection biases in richness-selected galaxy clusters for Euclid data.

Vera: Exactly, and what struck me most is how rigorously they handled the statistical modeling of those mass–richness relations using the Gamma-Poisson framework. It shows a serious commitment to making sure our final cosmological constraints are solid.

Subrahmanyan: From a theoretical standpoint, this calibration is vital because it directly addresses how observational selection effects can mimic or mask real cosmological signals in the mass–richness relation we’re trying to measure.

Jocelyn: It really does, and I think the use of an emulator trained on that data makes the whole process scalable for future surveys like Euclid. It’s a practical step toward reliable measurements.

Vera: That scalability is what excites me about it; if this method works robustly across different survey areas, we can be much more confident in the systematic errors we are seeing.

Subrahmanyan: The impact here is significant because it tightens the connection between observable galaxy cluster counts and the underlying dark matter distribution across cosmic time.

Jocelyn: It’s a huge step forward for using these rich samples as probes of large-scale structure evolution. It moves us closer to clean cosmological measurements with less uncertainty from instrumental artifacts.

Vera: I agree; it takes a massive amount of painstaking work to get that level of calibration, and this paper shows that hard work paying off in a very concrete way for Euclid planning.

Subrahmanyan: The next steps, as the authors suggested, using Graph Neural Networks and Bayesian Deep Learning are where we can really push the limits of how much physics we can extract from these cluster populations.

Jocelyn: I’m looking forward to seeing those next iterations; it sounds like they’re building a really comprehensive inference engine now. It keeps getting more complex, which is exactly what we need for this kind of science.

Vera: I'm genuinely excited to see what those advanced architectures can uncover in the data. This paper lays an excellent foundation for how we move from just observing data to truly understanding the underlying cosmic web structure.

Subrahmanyan: Indeed, this work on "Euclid preparation: CosmoPostProcess" provides a crucial methodological tool that allows us to tackle selection biases head-on, which is essential for unlocking the full potential of Euclid’s weak lensing measurements.

Jocelyn: It certainly sets a high bar for how we should approach these large-scale structure surveys going forward. We’ll be watching what they do next with those improved models.

R. Ingrao et al.

Euclid Collaboration

astro-ph.CO, astro-ph.IM

Submitted: 2026-05-04

Updated: 2026-08-25

Comments: Accepted for publication in Astronomy & Astrophysics (A&A)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: The paper details advanced methodological frameworks—including sophisticated statistical calibration and machine learning emulators—designed to accurately quantify and correct for selection bias

Key concepts

CosmoPostProcess
A simulation-based forward model algorithm calibrated to reproduce optical cluster observables exactly as measured by the Euclid space telescope. Its main deliverable is a correction for stacked surface-density profiles binned in richness and redshift, accounting for selection-related systematic effects.
Selection Bias Correction
The process of correcting measured stacked cluster profiles to reflect what they would look like if an unbiased reference sample were used. This correction models how the weak lensing signal is modified when using richness-selected samples identified in the photometric Euclid survey compared to an unbiased sample.
Graph Neural Networks (GNNs)
An advanced AI technique suggested for improving membership emulation by moving beyond standard fully connected networks. GATs are proposed because they can model spatial relationships in data structure better, allowing different weights to be assigned to neighbors based on their physical relevance.
Normalizing Flows
A statistical modeling approach proposed to model the full conditional probability mass function directly instead of using fixed parametric forms for likelihood estimation. This allows the model to learn the shape of complex distributions flexibly, rather than relying on strong assumptions.

Terminology

Summary

The paper details advanced methodological frameworks—including sophisticated statistical calibration and machine learning emulators—designed to accurately quantify and correct for selection bias in richness-selected galaxy clusters. This capability is critical for Euclid preparation, ensuring that measurements of the mass–richness relation are robustly calibrated against observational biases, thereby allowing reliable cosmological interpretation from large-scale surveys.

HOD Calibration: Modeling Mass–Richness Relations

The study establishes a rigorous statistical framework to model the mean relation between richness (N), mass (lambdā), and redshift. The calibration set utilizes an eight-tile subsample covering a total of 400 deg2 of the Flagship2 5200 deg2 catalogue, applying a conservative mass threshold of 10 13 h-1 M and fixing the pivot redshift to z piv = 1.25. To account for extra dispersion beyond simple counting noise, the authors adopt a Gamma-Poisson (negative binomial) model:

  • The latent rate is drawn from a Gamma distribution: about Gamma(kappa, theta).

  • The observed richness N is conditional on: N about Poisson.

This formulation leads to the negative binomial probability mass function for richness, P(lambda, N), which allows the calculation of the total log-likelihood over the calibration set (L). The resulting posterior constraints for key parameters—such as 10(M 1, sat) and alpha (the slope of the satellite term)—are presented in Table C.1, providing precise measurements like 12.19+0.008-0.008 for the characteristic mass scale and-0.208+0.003-0.013 for the index of redshift evolution (epsilon).

Membership Emulator Training and Inputs

A key component of the analysis is the training of a compact fully connected emulator for Pmem. This machine learning model predicts cluster membership probability using four primary inputs: z, r/r cl, z cl, and f bkg. The architecture is designed with three dense layers incorporating Leaky ReLU and dropout, culminating in a sigmoid output. Hyperparameters were optimized using Optuna, which allowed the authors to interpret the model's sensitivity via functional ANOVA.

The resulting analysis of input relevance confirms that certain variables are dominant predictors:

  • Fig. C.1 identifies z, R/R cl, and z cl as the leading predictors.

  • Crucially, the results show that f bkg contributing little, which supports the use of a simplified proxy for this background term while maintaining accuracy.

Validation and Consistency Checks

The emulator's performance was rigorously validated across multiple fronts. Figure C.2 demonstrates that the emulator successfully reproduces the dependence of Pmem on z and projected separation for an individual object, and, at the catalogue level, preserves the mass–richness relation with residuals consistent with zero within the uncertainty band.

Furthermore, a final consistency test compares selection bias calculated using CosmoPostProcess on stacked profiles against those derived from independent runs (RICH-CL and COMB-CL) on Flagship2. This comparison confirms that the framework is robust, as shown in Figure C.3, which compares the selection bias across different methodologies, ensuring that the final catalogue used for analysis remains accurate regardless of the specific processing pipeline employed.

Improvements for AI systems

This document describes sophisticated applications of machine learning (ML) and statistical modeling to astrophysical data (galaxy clusters). The current methods—using Fully Connected Networks (FCNs) for emulators and complex statistical fitting for HOD calibration—are state-of-the-art but can be significantly enhanced using modern, advanced AI techniques.

Here are the specific improvements I recommend, categorized by the system component, followed by what the improved AI system will achieve.


The current emulator uses a standard FCN architecture (Dense layers with Leaky ReLU and Dropout). This approach treats the input features (z, r/r cl, z cl, f bkg) independently. Since astrophysical data often possesses inherent spatial or directional dependencies, this is a limitation.

1. Architectural Improvement: Transition to Graph Neural Networks (GNNs)

  • Action: Instead of treating the input features as flat vectors for a standard FCN, model the cluster structure (or the local environment) as a graph. The nodes would represent individual objects/features (z cl, r/r cl, etc.), and the edges would encode their physical relationships (e.g., proximity in 3D space, or correlation based on angular separation).

  • Specific Technique: Implement a Graph Attention Network (GAT). GATs are superior to standard CNNs for this task because they assign variable weights (attention scores) to different neighbors, effectively determining which physical inputs are most critical at any given location.

2. Data Representation Improvement: Spatio-Temporal Feature Encoding

  • Action: Given that the input includes projected separation (r/r cl) and redshift (z), incorporate a specialized encoding layer that explicitly handles the geometric constraints of the survey volume.

  • Specific Technique: Use Spherical Harmonics or Wavelet Transforms on the angular coordinates (if available, or implied by f bkg). This captures large-scale cosmic structure information more efficiently than simple linear input features, making the emulator robust to varying field geometries.

The current HOD calibration uses a Negative Binomial model derived from a Gamma-Poisson process, which is statistically robust but relies on fixed functional forms and assumes independence between the latent rate and other parameters.

3. Inference Improvement: Bayesian Deep Learning for Parameter Estimation

  • Action: Replace the reliance on nested sampling (Nautilus) with a Bayesian Deep Learning (BDL) framework (e.g., using Variational Autoencoders or Monte Carlo Dropout).

  • Specific Technique: Train a BDL model to map the observed catalogue features (z, r/r cl, etc.) directly to the posterior distribution parameters (alpha, epsilon, sigma intr, M 1, sat, M min, sat). This allows the system to not just find a point estimate (the mean relation) but to provide a full predictive uncertainty envelope for every single cluster detection simultaneously.

4. Model Complexity Improvement: Non-Parametric Likelihood Estimation

  • Action: Instead of fixing the likelihood function (Negative Binomial, Equation C.3), use a machine learning approach that learns the full conditional probability mass function P(N lambda, M, z) directly from the data.

  • Specific Technique: Implement Normalizing Flows. Normalizing Flows allow us to model highly complex and multimodal distributions (like those governing cluster richness) by transforming a simple known distribution (e.g., Gaussian) through a series of invertible functions. This avoids making strong parametric assumptions about the underlying physical process, vastly improving robustness when data deviates from the idealized Negative Binomial form.

By integrating these advanced components, the resulting AI system moves beyond being a simple emulator or a fitting tool; it becomes a Unified Astrophysical Inference Engine.

  1. High-Fidelity, Uncertainty-Quantified Prediction: The system can predict P mem not just with an average value, but with a full probability distribution (via GNNs and BDL), quantifying the inherent physical scatter and systematic error for every single prediction point.

  2. Self-Correction and Anomaly Detection: By modeling the dependencies (GNN) and the underlying distribution (Normalizing Flows), the system can flag clusters whose properties fall outside the physically plausible manifold defined by P mem or HOD relations. This is critical for identifying potential instrumental artifacts or rare, unmodeled astrophysical populations.

  3. Accelerated Parameter Space Exploration: The BDL approach drastically reduces the computational cost of marginalizing over complex parameter spaces compared to nested sampling, allowing researchers to test dozens of alternative physical models (e.g., varying the functional form of alpha or epsilon) in minutes rather than days.

  4. Interpretability and Causality Mapping: The use of Graph Attention Networks (GATs) provides a built-in mechanism for feature attribution. The system can output why it made a specific prediction by highlighting which input features (z, r/r cl, etc.) contributed the most attention weight to the final P mem value, providing crucial physical insight into the selection bias and cluster formation mechanisms.

Abstract

We present CosmoPostProcess, a simulation-based forward-modelling algorithm calibrated to reproduce Euclid optical cluster observables. Its main deliverable is a correction for stacked surface-density profiles, binned in richness and redshift, accounting for selection systematics in richness-selected samples relative to unbiased references. We focus on the Euclid richness definition foreseen for cosmological analyses, which does not apply a colour selection; red-sequence richness is not considered. The algorithm processes N-body simulations by painting galaxies with a halo-occupation model and emulating survey detection and richness assignment. We also implement a novel estimate of optical cluster centres from projected galaxy densities, validated against Euclid pipelines. Baryonic effects are included through a correction calibrated on hydrodynamical simulations; the baryon-corrected excess surface density agrees within 2,% over r in[0.1,,5],h-1, Mpc. Selection-bias contributions are assessed by varying cosmology and the mass--richness relation. Projection-induced selection bias follows a robust pattern: correlated large-scale structure projected along the line of sight enhances the stacked profile near the one-halo to two-halo transition, peaking at about 1,h-1, Mpc with an amplitude of 20!-!40,%, depending on richness and redshift. The effect is mild at low and intermediate redshift (z 0.7), at the few-percent level, but becomes more relevant at higher redshift (z 0.7). Baryonic modifications remain sub-dominant outside the core, at about 2,% beyond r 0.3,h-1, Mpc. The framework delivers radial profile corrections with uncertainties, combining projection-induced selection bias, baryonic physics, and miscentring, to control systematics in Euclid DR1 cluster cosmology. (abridged)

Sources

Related papers