Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions

arXiv:2507.03074 · astro-ph.CO, astro-ph.GA · Submitted 2025-07-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: Next we'll be talking about the paper "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions".

Jocelyn: The paper was written by the authors from.

Vera: Stay tuned as we take you through the paper and discuss its implications.

Jocelyn: We also have Subrahmanyan with us today — guest researcher.

Vera: Alright, let's get started.

Summary: Vera: We've seen that this paper introduces a way to map initial conditions to final halo mass, so let’s look closer at what this mapping actually is. The authors describe a U-Net architecture that achieves this field-level prediction.

Jocelyn: So, the AI isn't just looking at a few particles in a patch; it’s taking the entire three-dimensional density field as input, which is impressive for an AI model.

Subrahmanyan: Yes, and because the initial density field contains all that information about where collapse will happen, Subrahmanyan thinks this U-Net is perfectly suited to identify those protohalo patches before they become fully virialized structures.

Vera: They train it using two thousand independent fast N-body simulations, which are run using the Disco-Dj code, making sure the training data covers a wide range of cosmological parameters.

Jocelyn: That vast dataset is key for "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions," ensuring that our AI isn't just trained on one specific version of the universe.

Subrahmanyan: Subrahmanyan views this comprehensive training as necessary to make sure the model handles all those varied cosmological settings robustly, which is far beyond what older fixed models could do.

Vera: The paper’s summary shows that this approach is designed to bypass the traditional non-linear N-body simulation in certain analysis steps by directly utilizing the linear density field, giving us a very efficient pathway to results.

Jocelyn: Efficiency is important when we have so much data to process, especially when running large-scale simulations like those used for our galaxy surveys.

Subrahmanyan: And Subrahmanyan notes that this framework allows us to treat the initial conditions as a latent variable, which is a massive conceptual leap in how we approach cosmological inference.

Vera: It’s all about making the the mapping between the initial field and "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions" much more direct than previously possible.

Jocelyn: This opens up so many new avenues for us to explore how structure forms across cosmic time, especially when we can see where it starts.

Subrahmanyan: Indeed, Subrahmanyan believes this is a foundational tool for the next generation of cosmological simulations and analysis.

Improvements: Vera: Moving on to the results, we've seen how the model works; now let’s talk about its performance in "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions." The authors found that their U-Net was quite conservative.

Jocelyn: That means it tended to be cautious, right? It didn't predict as many haloes as were actually found in the simulations, which is an important nuance for us to understand.

Subrahmanyan: Yes, Subrahmanyan points out that this conservatism is partly due to the inherent difficulty of defining boundaries in a smooth field and that we are still working with discrete particle assignments.

Vera: Despite being conservative, it's very accurate on average, achieving an error level of about eight percent, which is competitive with established methods.

Jocelyn: That’s great news for the practical application, knowing that we can trust the predictions within that margin of error for our surveys.

Subrahmanyan: Subrahmanyan also highlights how they used a specific loss function—the squared Earth Mover’s Distance or EMD2—to make sure the model didn't try to "interpolate" through low-mass bins.

Vera: That means the AI is forced to be physically consistent and not just guess that a particle in a patch belongs to a small, weak halo when it should be part of a massive one.

Jocelyn: It prevents those spurious low-mass predictions, which is something we often worry about when trying to classify objects in our own observational data.

Subrahmanyan: Subrahmanyan believes this loss function is key to maintaining the physical integrity of the "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions" output across mass scales.

Vera: We’ve also seen how well it performs at identifying background versus halo, with a very high true negative rate—around ninety-five point six percent.

Jocelyn: That high accuracy in identifying what's *not* a halo is crucial for filtering out the noise in our large datasets, making the whole process much cleaner.

Subrahmanyan: And Subrahmanyan notes that this performance, combined with its ability to capture the parameter dependence, represents a significant leap forward compared to previous methods that are now struggling with cosmic variance.

Vera: It' seems like a real improvement in robustness across different simulation realizations.

Jocelyn: We’re ready to discuss how this consistency translates into big picture scientific discoveries, though, which Subrahmanyan is also keen to hear about.

Implications: Vera: Now we are looking at the big implications of "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions." The core capability here is that since the U-Net is differentiable, we can calculate gradients.

Jocelyn: So, instead of just running two separate simulations to see how a change in cosmology affects our results, we can actually take the derivative of the predicted HMF with respect to the cosmological parameters.

Subrahmanyan: Precisely, Jocelyn. Subrahmanyan thinks this is revolutionary for parameter inference because it allows us to use gradient-based methods like Hamiltonian Monte Carlo or HMC to find the best-fit cosmology much more efficiently than before.

Vera: We can actually calculate how the HMF changes when we vary parameters like m or n s, which is a huge deal for characterizing our universe.

Jocelyn: It’s not just about finding one specific answer; it' about understanding the sensitivity, so seeing where the HMF is most sensitive to tells us what aspects of cosmology are most important to look at.

Subrahmanyan: Subrahmanyan notes that we can even use these gradients to extrapolate existing models, which allows us to test scenarios outside of the original parameter space used for training.

Vera: That ability to push beyond the original range is powerful, allowing us to investigate areas like how a new form of dark energy might affect halo formation.

Jocelyn: It also lets us explore what happens if we change the fundamental properties of our initial density fluctuations, which is something we can't easily do in traditional simulations.

Subrahmanyan: Subrahmanyan agrees that this opens up complex, multi-dimensional studies on the nature of dark matter and its interaction with gravity.

Vera: We can even see how the HMF responds to changes in the white noise amplitude of the initial conditions, which is a unique way to study cosmic variance itself.

Jocelyn: It’s like we're measuring how much noise in our starting point affects the final structure, which is really important for us as observers trying to account for uncertainties.

Subrahmanyan: Subrahmanyan sees this as an essential tool, allowing us to directly quantify the impact of the initial conditions on "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions."

Vera: It’s a framework that allows us to explore how we can use machine learning not just for prediction, but for sophisticated scientific inference.

Jocelyn: We're really excited about the possibilities this opens up for future research, too.

Conclusion: Vera: So, we've covered a lot of ground on "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions," from how it works to its potential.

Jocelyn: We can say that this model is a very robust way to predict halo mass while simultaneously capturing how cosmological changes affect those predictions.

Subrahmanyan: Subrahmanyan is confident that this technique will become a standard tool for simulating the universe in a way that is both physically grounded and mathematically tractable.

Vera: It’s great that the authors are suggesting continuous future work, like moving toward predicting continuous halo properties rather than just binned ones.

Jocelyn: And I think it’s also interesting that we can use these derivatives to correct other models, which is a massive practical application of this whole framework.

Subrahmanyan: Subrahmanyan emphasizes that this ability to capture the subtle trends in the HMF is what makes "Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions" so valuable for connecting theory and empirical evidence.

Vera: It's been a really informative discussion, and I think we’re all incredibly excited about what’s to come with this research.

Jocelyn: We are, Vera; it feels like a massive step forward in the future of data-driven astronomy.

Subrahmanyan: And Subrahmanyan hopes that this is just the first of many breakthroughs utilizing this technique, setting the stage for even more complex cosmological studies.

astro-ph.CO, astro-ph.GA

Submitted: 2025-07-03

Updated: 2026-08-23

Comments: 14 + 5 pages, 11 + 5 figures, Published in the Open Journal of Astrophysics, comments welcome

Journal ref: Open Journal of Astrophysics, Vol. 9, 2026

DOI: 10.33232/001c.168410

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: " * Motivation and Background Modern cosmological inference requires models that are "differentiable to enable efficient, gradient-based parameter estimation and uncertainty quantification." The halo

Key concepts

U-Net architecture
This is an AI model used to map initial conditions, specifically a three-dimensional density field, directly to final halo mass. It is suited for identifying protohalo patches within the initial density field before they fully virialize.
Differentiable Halo Mass Prediction
Because the U-Net model is differentiable, researchers can calculate gradients of the predicted halo mass function with respect to cosmological parameters. This allows for efficient gradient-based methods like Hamiltonian Monte Carlo to find best-fit cosmologies.
Earth Mover’s Distance (EMD2)
This specific loss function was used during training to ensure the AI model maintains physical consistency across different mass scales. It prevents the model from incorrectly interpolating through low-mass bins, forcing it to be physically accurate.
Cosmological Inference
The framework allows researchers to use gradients of the halo mass function to efficiently perform parameter inference. This lets them understand which cosmological aspects are most important and enables extrapolation beyond the original training parameter space.

Terminology

Summary

"


Motivation and Background

Modern cosmological inference requires models that are differentiable to enable efficient, gradient-based parameter estimation and uncertainty quantification. The halo mass function (HMF) is a sensitive probe of cosmology, yet traditional methods for predicting its form—such as the excursion set formalism or the peak-patch model—are often discrete and lack the necessary differentiability. This paper addresses this gap by presenting a novel approach for predict[ing] the abundance of dark matter haloes and their cosmology dependence using a differentiable, field-level neural network (NN) model.

Methodology: Data Generation

The study utilized fast N-body simulations generated using the Disco-Dj code. A large training/validation dataset consisting of 2000 independent simulations was created. The cosmological parameters varied across a three-dimensional cube, including the matter density parameter m, the linear power spectrum normalization sigma 8, and the primordial scalar spectral index n s.

Methodology: NN Architecture and Training

The core task is to assign a halo mass to every Lagrangian fluid element (voxel) from the initial linear density field. The model employed is a 3D U-Net architecture, which was chosen because its structure bears a fundamental similarity to what convolutional neural networks (CNN[s]) should be able to capture.

The network was trained as a multiclass semantic classification task, where each Lagrangian fluid element is classified into one of 12 bins:

  • The inner 10 bins contain masses between 4 times 10 12 M/h and 1 times 10 15 M/h with logarithmic spacing.

  • Bin 0 represents the background (non-halo) elements.

  • The uppermost bin collects all haloes above 1 times 10 15 M/h.

The training process involved generating 96,000 data cubes through rotation and flipping (data augmentation). The network was trained using Adam gradient descent with the squared Earth Mover’s Distance (EMD squared) loss. This specific loss function is crucial because it measures the minimum work required for transforming one probability distribution into another and does not reward assigning probability mass to intermediate bins unless they are actually supported by the data, thereby preventing spurious interpolation through low-mass bins at halo edges.

Results: Classification Performance

The classification performance of the U-Net was evaluated against established methods (e.g., López-Cano et al. 2024). The model demonstrated a high True Negative Rate (TNR = 95.6%) and an overall accuracy of 90.7%. However, the authors note that the U-Net is rather conservative in predicting protohaloes, achieving only 73.2% True Positive Rate (TPR).

Results: Halo Mass Function (HMF) Prediction

The HMFM is computed directly from the U-Net output by summing class probabilities over the entire box. This allows for a differentiable calculation of the HMF without post-processing. In comparison to 100 different test simulations at fixed cosmological parameters, the U-Net's average prediction shows an accuracy of approximately 8%. The authors observe that while the model performs well for small and medium-sized halo masses, it underestimates higher mass haloes due to cosmic variance.

Results: Cosmology Dependence and Derivatives

A key feature of the the framework is its ability to compute gradients with respect to cosmological parameters. Using forward-mode differentiation via Jax, the authors calculated the dependence of the HMF on m, n s, and sigma 8.

  1. ** m (Matter Density):** The U-Net gradient matches quite well with analytical models and emulators.

  2. ** n s (Spectral Index):** The model captures the trend accurately, though a small discrepancy exists between the analytical and emulator HMFs.

  3. ** sigma 8 (Power Spectrum Normalization):** This parameter overestimates the dependence for lower mass bins compared to other methods, which is noted as being difficult to capture accurately due to its sensitivity.

Furthermore, the model's gradients were tested against 50 independent realizations of the initial density field. The results showed that large-scale fluctuations have the most significant impact on the HMF, with increasing white noise amplitude on large scales reducing the abundance of low-mass haloes, while enhancing the formation of intermediate-mass haloes.

Results: Extrapolation and Application

The paper demonstrated how to use differentiable mass functions to enhance non-differentiable HMF models. By applying a first-order Taylor expansion around a fiducial cosmology using gradients extracted from the Tinker08 HMF, the resulting approximation reproduces the perturbed Mira-Titan HMF at the sub-percent level. This capability allows for parameter inference even when a baseline model (like Mira-Titan) is not inherently differentiable.

Conclusion

The study confirms that automatically differentiable HMFs can be used as gradient information for parameter inference and capture cosmological dependence reasonably well. The proof-of-concept model was found to be accurate at the few-percent level, with future improvements suggested by moving beyond binned labels to continuous properties and utilizing a 'windowed labelling' approach.

Improvements for AI systems

The core improvement involves replacing or augmenting standard classification loss functions with Earth Mover's Distance (EMD) losses, specifically the squared EMD (EMD 2), when the underlying physical process is known to be bimodal or multimodal with significant gaps between plausible states.

Improvement: Integrate the expected EMD 2 loss, as defined by Equation (C1), into the training objective function (L Total = + lambda times E[LEMD 2]). This loss measures the discrepancy between the predicted probability distribution p and the expected ground truth distribution e, where e is a weighted mixture of one-hot target vectors corresponding only to plausible classes.

What the Improved AI System Can Do:

  • Enforce Sparsity of Prediction Support: The system will be rigorously trained to assign near-zero probability mass (p i about 0) to intermediate or unsupported classes (like the low-mass halo bin in the example) when the true physical state is confidently expected to belong only to two distinct, separated modes (e.g., background or high-mass halo).

  • Prevent Spurious Interpolation: Unlike standard cross-entropy (L X), which can penalize probabilities exponentially as they approach zero near unsupported regions, EMD 2 maintains a stable gradient structure that guides the network to maintain probability mass only on the two plausible extremes, preventing spurious assignments to intermediate bins.

  • Robustness in Physical Inference: In fields like astrophysics or materials science where observations yield clear separation between regimes (e.g., diffuse vs. dense), this system will provide a more physically constrained inference, ensuring that the predicted state p accurately reflects the known physical manifold of the data, rather than interpolating along a mathematically smooth but physically invalid path.

  • Context-Aware Uncertainty Quantification: The system can quantify uncertainty by calculating the minimum EMD 2 loss required to satisfy the constraints of known physical regimes. If the network's prediction p deviates significantly from the optimal manifold (e.g., by assigning weight to a state outside the expected [alpha, 0, 1-alpha] configuration), a large penalty is immediately incurred.

  • Improved Decision Boundary Sharpness: The training process will lead to sharper, more reliable decision boundaries that respect known physical gaps between classes, leading to higher fidelity segmentation or classification outputs compared to models trained with standard losses that favor smooth transitions everywhere.

Feature Standard Cross-Entropy (L X) EMD 2 Loss (Proposed) Benefit/Capability Gain

:---:---:---:---

Handling Gaps (Unsupported intermediate states) Tries to smoothly interpolate, potentially penalizing zero probability incorrectly. Diverges near true class boundaries if prediction is wrong. Explicitly minimizes loss by assigning zero weight to unsupported bins, focusing mass only on plausible extremes. Physical Fidelity: Prevents the model from hallucinating or spuriously assigning intermediate states.

Loss Landscape (General) Can be sensitive to local minima and exhibit complex geometry near edges. Convexity rewards hedging between plausible outcomes; the loss landscape guides optimal sparsity. Training Stability: More robust training, especially when the true distribution is highly bimodal or sparse in feature space.

**Output p ** (Optimal) May distribute probability mass across irrelevant bins if not strictly constrained. Optimal prediction p concentrates mass only on the most likely plausible classes (p = (alpha, 0, 1-alpha)). Interpretability & Reliability: The output directly reflects the known physics of the system, providing a highly reliable estimate of multimodal probability.

Abstract

Modern cosmological inference increasingly relies on differentiable models to enable efficient, gradient-based parameter estimation and uncertainty quantification. Here, we present a novel approach for predicting the abundance of dark matter haloes and their cosmology dependence using a differentiable, field-level neural network (NN) model, and study how well the cosmology dependence is captured by common parametrisations of the halo mass function (HMF), and by our NN-based approach. By training a 3D U-Net on initial density fields from fast N-body simulations with varying cosmological parameters, we enable direct, differentiable mapping from the linear density field to protohalo patches and their mass bins. Our method achieves competitive accuracy in identifying protohalo regions and in capturing the dependence of the HMF on cosmological parameters. Our NN derivatives agree well with finite differences of both analytical and emulated HMFs, at the level of the disagreement among the different models. We further demonstrate how the NN model can additionally be used to investigate the response of the HMF to changes in the initial Gaussian random field. Finally, we also demonstrate that a differentiable model can be used to extrapolate existing models at very high precision.

Sources

Related papers