Fine-tuned Normalizing Flows for ALICE Zero Degree Calorimeter Fast Simulation

arXiv:2608.12795 · physics.ins-det, cs.LG, hep-ex · Submitted 2026-08-13 · Read on arXiv

Emilia Majerz, Jacek Otwinowski, Witold Dzwinel, Jacek Kitowski

AGH University of Krakow · The Henryk Niewodniczański Institute of Nuclear Physics, Polish Academy of Sciences

physics.ins-det, cs.LG, hep-ex

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: This paper has been accepted for presentation at the 16th International Conference on Parallel Processing & Applied Mathematics (PPAM 2026)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: Simulating the ALICE Zero Degree Calorimeter (ZDC) neutron detector responses at the LHC is computationally expensive, requiring complex Monte Carlo chains.

Terminology

Summary

Simulating the ALICE Zero Degree Calorimeter (ZDC) neutron detector responses at the LHC is computationally expensive, requiring complex Monte Carlo chains. We develop a generative surrogate, focusing on Normalizing Flows (NFs). Through transfer learning, we pretrain on the full imbalanced dataset and fine-tune specialized models for different particle types (γ, n, Λ, KS0, Σ +) using two gradual-unfreezing schemes. As standard ZDC metrics like Wasserstein distance overlook conditional structure, we introduce refined metrics: conditional weighted MAE, dispersion ratio, and Jaccard co-activation error, that better capture physics-relevant input-output dependencies and response variability. Our ensemble of fine-tuned models achieves a Wasserstein distance of 1.61 ± 0.02, outperforming baselines across all metrics. This work provides a generalizable NF-based framework for LHC detector simulation, combining NFs, conditional fine-tuning, and physics-motivated evaluation.

The dataset used in this study contains N=306,780 samples generated by the GEANT4 software. Each sample is a pair consisting of a vector with features of particles produced by primary proton-proton collisions (called primary particles) – such as energy, three-momentum, position in three dimensions, mass, and charge – and a 44 × 44 matrix representing the ZN response. Each matrix element represents the number of photons detected in the corresponding fibre. The detector response is therefore treated as a single-channel image, where each pixel value corresponds to a photon count. To suppress low-signal random responses, we retain only matrices with at least 10 photons; after this filtering, the median photon count is 234.

An important property of the dataset is its strong class imbalance. The γ particles account for almost 61% of the data, n for over 23%, p for over 5%, Λ for over 3%, and the remaining 17 particle types together for around 7.5%. This imbalance may bias the model toward the dominant particle types and degrade performance on underrepresented ones. Possible remedies include rebalancing the dataset, applying standard techniques for imbalanced data, or training dedicated models for poorly represented particle types. Here, we investigate the last strategy.

Among the particles used in our experiments, the fractions of primary particles that hit the detector directly and produce a response of at least 10 photons are 45% for γ, 78% for n, 46% for Λ, 21% for KS0, and 0% for Σ +. When the fractions are lower, we should expect multiple showers and scattered responses induced by secondaries.

The diversity of detector responses is an additional challenge for any data-driven surrogate model. The same vector of input parameters can produce substantially different ZN responses. This happens because particles traverse a long distance between IP2 and the detector, during which secondary particles may be produced and primary trajectories may deviate from their original paths. Different particle types also exhibit different levels of response diversity: some produce fairly consistent showers, while others lead to much more variable shower positions and shapes. To quantify this effect, we use the variance-based diversity measure introduced in [4]: fdiv (c) = P i,j q 1 Xc P t∈Xc xtij − µij 2, where xtij is the pixel value at coordinates i, j for the detector response t, µij is the mean pixel value at i, j for the unique input vector c, and Xc is the number of detector responses associated with input c. This quantity is positively correlated with particle mass and energy.

We use a three-stage approach: 1. Pre-train an NF model on the full, imbalanced dataset to capture general detector-response patterns. 2. Fine-tune the pre-trained model separately for selected particle types using the corresponding subsets. 3. Combine the fine-tuned models into an ensemble and verify its performance. We split the dataset into training, validation, and test sets in a 70:10:20 ratio. Models with the best validation performance were selected for testing. During fine-tuning, we used only particle-specific samples from the corresponding splits to avoid data leakage. Each model was evaluated five times, generating five samples per input vector on the test set.

In our initial experiments, we trained NFs directly on the primary particle features. To reduce the dynamic range of the detector-response images, we applied a log transform to the pixel values. Although the generated showers were visually plausible, the predicted photon counts did not match the data well. We therefore adopted a two-stage approach: the NF generates a normalized response shape, which is then rescaled by an auxiliary model predicting the expected total photon count. We also provided the expected photon count as an additional input feature during NF training, following [12]. We compared several models for predicting the photon sum, including NFs, SVMs, GBRs, and BNNs. The BNN performed best and was used in all experiments (MAE: 53.20, RMSE: 111.47). Although the absolute errors appear large, the median percentage error of the best BNN was 12%, which seems acceptable given the stochastic nature of the target.

For the NF model, we follow the configuration of [20], itself based on [12]. The model uses the MAF architecture with rational quadratic spline (RQS) transformations [6], whose parameters are produced by MADE blocks [9]. The NF maps the data distribution to a simple base distribution through a sequence of invertible transformations. Sampling proceeds in the reverse direction: a point is drawn from the base distribution, the MADE blocks compute the spline parameters, and the transformations are applied successively. Permutation layers are inserted between spline blocks to increase flexibility. Compared with [20], we improved the baseline by replacing the conditional diagonal Normal base with a standard Normal and using the pixel-value recalculation post-processing scheme.

Because multiple detector responses may be valid for the same input, direct pixel-wise comparison is not meaningful. Following prior work on the ZN detector [3, 5, 11, 20, 4], we use the mean 1-Wasserstein distance between the distributions of photons collected by each PM. We refer to these groups as channels and report the average over m = 5 channels: W1 (w, ŵ) = 1 m Pm i=1 R 1 0 Fw−1 i (z) − Fŵ−1 i (z) dz. Here, Fw−1 i is the inverse cumulative distribution function of channel distribution wi, and ŵi denotes the predicted distribution. We compute channel values by applying five checkerboard masks to each response image and summing the values within predefined fibre groups.

Although the Wasserstein distance captures global distributional similarity, it does not assess whether the response matches the input condition. Thus, we compute a conditional mean absolute error over unique inputs, as in [16]: MAEcond (w, ŵ) = P 1 k pk Pm i=1 P k pk · wki − ŵik. We use logarithmic weights, pk = log(1 + nk), to reduce the influence of highly repeated conditions while still favouring well-sampled regions of phase space.

To quantify channel variability, we use a per-channel median absolute deviation (MAD) with shrinkage toward a particle-level prior. MAD is robust to outliers, which is useful for detector responses that may contain rare, extreme showers, and shrinkage further stabilises it in low-sample cases. We compute the prior MADparticle as a weighted median MAD value computed for reference inputs with at least 5 responses available. For each unique input and channel: MADu,c = mediani Ru,c i − medianj (Ru,c j), MADg u,c = αu MADu,c + (1 − αu) MADparticle,c, αu = Nu Nu+κ. Here, κ is the 10th percentile of the number of responses per particle type. We compare reference and generated variability using the dispersion ratio: δu,c = MADg gen u,c − MADg ref u,c MADg gen u,c + MADg ref u,c + τ, which provides a bounded symmetric score, treating overdispersion and underdispersion equally. We compute this separately for each channel and report the mean over channels, followed by a weighted mean over unique inputs. Cases with fewer than two samples are excluded.

We also measure spatial co-activation between neighbouring channels using a Jaccard-based error. For each event, we define a binary hit indicator and compute the Jaccard index for each neighbouring pair (c1, c2): Jc1,c2 = P i 1(Bu,c1 i = 1 ∧ Bu,c2 i = 1) P i 1(Bu,c1 i = 1 ∨ Bu,c2 i = 1), Bu,c i = 1(Ru,c i > 0). If both channels are always zero, we set Jc1,c2 = 1. The final score for a unique input is JaccErru = 1 P P pairs Ju gen − Juref, where P is the number of neighbouring channel pairs. We report the weighted mean over unique inputs.

After training on the full dataset, we fine-tune separate models for selected particle types. We focus on particle types for which the baseline Wasserstein score is relatively poor and exclude very rare classes from the fine-tuning study. The selected particles are γ, n, Λ, KS0, and Σ +. Although fine-tuning all particle types might yield further gains, our goal here is to validate the methodology rather than to optimise every class exhaustively. Standard neural-network fine-tuning usually freezes early layers and retrains later ones. For NFs, we consider two gradual-unfreezing directions: from the base-distribution side toward the data side, and in the opposite direction. We evaluate both because the flow structure makes either choice plausible. In each step, we unfreeze one MADE block at a time.

Discrete-valued data, such as detector responses, are often modelled by adding small uniform noise so that each integer value corresponds to a small interval. We use the same idea here. In the two-stage setup, the NF first generates a normalized response shape, which is then rescaled by the predicted total photon count (including the noise contribution). After this rescaling, we remove the added noise using one of four schemes: 1. Apply the floor function. 2. Subtract the expected preprocessing noise and round to the nearest integer. 3. Apply the floor function and then renormalize and rescale the response to preserve the total photon count. Finish with rounding. 4. Subtract the expected preprocessing noise, round, and then renormalize and rescale. Finish with rounding. The renormalization steps are used because direct floor or round operations may force some pixels to zero and distort the total photon count. After renormalization, we rescale the responses, and round the values to integers again.

The fine-tuned models can be combined with the baseline model in an ensemble. For each input, we use the dedicated fine-tuned model if one exists; otherwise, we use the baseline model trained on the full dataset. The ensemble therefore assigns specialised models to particle types that benefit from fine-tuning while retaining a shared model for the rest. In principle, the submodels may also differ in architecture, but in this work we use a uniform MAF-based setup.

The chosen post-processing approach has a large effect on Wasserstein performance, and no single method is optimal across all particle types. Method 3 is best overall, but particle-specific optima differ. Using the default post-processing method, the Wasserstein score for the whole dataset is 1.73 ± 0.02. When a dedicated method is selected for each particle type, the score improves slightly to 1.70 ± 0.02.

We fine-tuned the baseline NF model (Baseline 100) for γ, n Λ, Ks0, and Σ+ particles, separately. The model was first trained for 100 epochs and then fine-tuned for another 100 epochs using the two schemas introduced in Section 3.3: FT 100 ND, where unfreezing starts from layers close to the Normal distribution, and FT 100 DN, where unfreezing starts from layers close to the data. We also compared the results with a baseline model trained for 200 epochs (Baseline 200). In addition, we trained separate models from scratch for each particle type (Individual), with 100 and 200 epochs, but these performed substantially worse than the other approaches.

We first assess the NF models responsible for modelling shower position and shape independently of the model predicting the expected number of photons. To do this, we use the photon counts derived from the dataset as input. The analysed particles differ substantially in sample size, in the average number of responses per unique input, and in their level of diversity. None of the approaches is best in every case, but FT 100 ND stands out as the strongest overall. The Σ + case remains exploratory because only a small number of samples is available. With that caveat, FT 100 ND performs best in 3 out of 4 cases for the MAEcond, δ, and JaccErr metrics. These metrics are computed at the level of unique input vectors, preserving the physical input-output relationship. By contrast, the Wasserstein score is more global and can favour distributional similarity even when the conditional structure is less accurate. FT 100 DN performs better than the baseline in all metrics for the γ case, which suggests that this fine-tuning direction can also be useful for some datasets.

To assess whether the differences between models are statistically significant, we performed bootstrap resampling of the unique-input level for the metrics MAEcond, δ, and JaccErr. We focus on FT 100 ND because it usually performs better than FT 100 DN, and on the Baseline 200. We exclude Σ+ due to small number of samples available. First, we computed each metric at the unique-input level and averaged the results over five runs. Then, for each of the 10,000 bootstrap samples, we resampled unique inputs, preserving all associated responses and weights, and computed the differences between the fine-tuned and baseline models. This yielded p-values, 95% confidence intervals, and the fraction of bootstrap samples favouring the fine-tuned model. For the γ particles, all fine-tuned results are significantly better than the baselines. This suggests that, when enough particle-specific samples are available, fine-tuning is the most efficient way to obtain a surrogate model. For the other cases, the conclusions are less clear: some results are statistically significant in favour of the fine-tuned models, while none favour the baseline. Whenever the fine-tuned models outperform the baseline in Table 3, the majority of bootstrap samples follow the same trend. However, this analysis should be interpreted cautiously. In some cases, only a small number of samples is available per unique input. In addition, resampling the observed unique inputs changes the effective input distribution, which may not correspond to a physically plausible perturbation.

Here, we analyse the performance of the full two-stage pipeline. For each particle and setup, we report scores obtained using the best-performing noise-removal method with respect to the Wasserstein score. Because training models from scratch for each particle performs significantly worse than other approaches, it is excluded from the analysis. Fine-tuning outperforms all other approaches for γ, Λ, and Σ + in terms of Wasserstein score. For KS0 and n, fine-tuning results are only slightly worse than the baseline. The best noise-removal method is consistent between baseline and fine-tuned models for each particle. The only exception is n, where floor-based methods with and without recalculation perform best in different settings. This suggests that the model behaviour is consistent and that a once-chosen noise-removal method for a given particle should also work well for similar models. Compared to Table 3, the full pipeline shows worse metric values, especially for Λ, KS0, and Σ +. This indicates that the auxiliary photon-count predictor is the main bottleneck, as it fails to learn the input-to-photon mapping well for these particles.

The other metrics also generally favour fine-tuning. Interestingly, FT 100 DN sometimes performs best, unlike in the previous NF-only analysis. The metric that remains most consistent between Tables 3 and 5 is Jaccard index error; the only significant change in the best-performing model is for the Λ case, which is not surprising given the much worse Wasserstein score observed there. The most stable results are again obtained for γ, which is the least diverse particle type and also the most abundant in the data set. Here, FT 100 ND outperforms the rest. For Σ +, FT 100 ND is the best in 3 of the 4 metrics and second best in the remaining one, which also makes it stand out. For the remaining cases, the ranking is less straightforward. We therefore rank the results within each metric and combine them into a single averaged value. For n, Baseline 200 and FT 100 DN perform similarly, although Baseline 200 is only slightly better in Jaccard index error, which suggests that the FT 100 DN can be considered to be better overall. For Λ, FT 100 ND and FT 100 DN also have the same mean ranking, and the slight advantage of FT 100 DN becomes visible only when Table 3 is taken into account. For KS0, Baseline 100 and FT 100 ND have the same ranking, but the fine-tuned model is preferable when Table 3 is considered as well. Across the four better-sampled particle types (γ, n, Λ, KS0) and the four metrics, the fine-tuned FT 100 ND model is best overall in most comparisons. The difference between results presented in Tables 3 and 5 suggests that the main bottleneck in this study is the prediction of the expected photon number, so further work should focus on that stage.

We also want to highlight that fine-tuning is more efficient than training the baseline longer on the full dataset. Thus, in cases where the fine-tuning improves the Baseline 200 or where it behaves similarly, it makes more sense to follow the fine-tuning approach, which saves time and reduces energy consumption.

We construct the final ensemble by selecting the best-performing model for each particle: FT 100 ND for γ, KS0, and Σ +; FT 100 DN for n and Λ; and Baseline 100 for the remaining particles (six models total). The ensemble outperforms both baselines, validating the multi-model approach. The ensemble achieves a Wasserstein distance of 1.61 ± 0.02, MAEcond of 6.29 ± 0.02, δ of 0.163 ± 0.001, and JaccErr of 0.121 ± 0.001, outperforming both Baseline 100 (WS 1.70 ± 0.02, MAEcond 6.57 ± 0.04, δ 0.176 ± 0.001, JaccErr 0.127 ± 0.000) and Baseline 200 (WS 1.76 ± 0.02, MAEcond 6.43 ± 0.03, δ 0.172 ± 0.001, JaccErr 0.125 ± 0.001).

MAF-based NF architectures are fast in evaluating the likelihood, which allows for reasonably short training, and slow during sampling. Our models require around 160.0 ms for response generation, compared to e.g. 0.023 ms needed for a GAN (NVIDIA A100 40 GB GPU, batch size 256) [20]. However, it is possible to train an Inverse Autoregressive Flow (IAF) student, which mimics the behaviour of the previously trained MAF teacher. This way, we can expect to obtain around 420 times faster models (0.38 ms per sample), as in [16].

Good performance alone is insufficient; models should also exhibit physically plausible reasoning. We use SHAP analysis to verify the ensemble’s feature dependencies for three quantities: total photon count, and shower centre x and y coordinates. For photon count, we explain the BNN predictions directly. For coordinates, we train surrogate BNNs on NF-generated data through knowledge distillation [10], as a direct NF explanation is too slow. The analysis indicates that the photon count depends primarily on Energy and momentum in the z direction Pz, with higher values producing more photons, as expected physically. Because we removed the sign from Pz (indicating ZDC side, negligible here), its contribution is similar to Energy due to the energy-momentum relation E 2 = Px2 + Py2 + Pz2 + mass2 (in the dataset, the Pz value is orders of magnitude bigger than Px, Py and mass, thus dominating the E value). Shower x and y coordinates depend mainly on Px, Py, and Energy, as expected from particle trajectory physics. The influence of charge reflects magnetic field deflection near the detector.

We applied Normalizing Flows (NFs) to simulate ZDC neutron calorimeter responses, demonstrating substantial improvements through transfer learning, fine-tuning, and ensembling. Our ensemble generator achieves a Wasserstein distance of 1.61 ± 0.02, outperforming all baselines across all metrics while maintaining physically plausible feature dependencies verified via SHAP analysis. Importantly, fine-tuning saves time and reduces energy consumption compared to a longer training on the full dataset. Standard ZDC evaluation metrics, such as Wasserstein distance and per-pixel MAE, have known limitations. We therefore introduce conditional metrics: weighted channel MAE, dispersion ratio, and Jaccard co-activation error, which better capture input-output physical consistency across diverse particle types and response variability. Key methodological advances include: 1. Two fine-tuning schemes using gradual unfreezing, with one of them showing the strongest overall performance. 2. Four post-processing methods for noise removal, where particle-specific optimization yields further gains. 3. An ensemble combining specialised models that validates the multi-model approach. The analysis reveals that photon-count prediction remains the primary bottleneck, suggesting clear directions for future improvement. While current inference is slower than some other generative alternatives, IAF distillation offers a path to a substantial speedup. We also highlight the possibility of utilising different approaches to ensemble modelling, e.g., with a meta-model designed to assign a simulation task to a particular generative surrogate. This work provides a generalizable procedure for constructing a data-driven surrogate model of a complex numerical simulation, specifically aimed at predicting detector responses to particle parameters recorded at the LHC. The combination of normalizing flows, conditional fine-tuning, and physics-motivated evaluation establishes a strong foundation for production-ready fast simulation.

Improvements for AI systems

Based on the paper, here are the specific improvements you can make to AI systems and what the improved system can do:

1. Implement transfer learning with gradual unfreezing for generative models

  • Pre-train a normalizing flow on a large, imbalanced dataset, then fine-tune specialized models for underrepresented classes using two directional unfreezing schemes (from base distribution to data, and vice versa).

  • The improved system can generate high-quality synthetic detector responses for rare particle types (e.g., Σ+, KS0) without requiring large per-class datasets, reducing training time and energy consumption by 50% compared to training from scratch or longer baseline training.

2. Add conditional evaluation metrics for physics-aware validation

  • Replace global metrics (e.g., Wasserstein distance) with three new metrics: conditional weighted MAE (measures input-output dependency accuracy), dispersion ratio (quantifies response variability match), and Jaccard co-activation error (measures spatial correlation between neighboring channels).

  • The improved system can detect when a generative model produces statistically similar outputs but fails to preserve physical input-output relationships, enabling more reliable deployment in scientific simulation tasks where conditional accuracy matters.

3. Use a two-stage generation pipeline with auxiliary predictor

  • Decouple shape generation (normalizing flow) from total intensity prediction (Bayesian neural network), where the NF generates a normalized response shape and the BNN predicts the total photon count.

  • The improved system can handle stochastic, multi-modal outputs (where the same input yields multiple valid responses) with a 12% median percentage error in total count prediction, avoiding the failure mode of direct pixel-wise generation.

4. Apply post-processing noise-removal schemes with renormalization

  • Use four distinct noise-removal methods (floor, round, floor+renormalize, round+renormalize) after adding uniform noise to discrete data, with particle-specific optimal selection.

  • The improved system can generate integer-valued outputs that preserve total photon counts and spatial structure, reducing Wasserstein distance by 0.03–0.04 compared to naive rounding, and can automatically select the best method per data class.

5. Build an ensemble of specialized generative models

  • Combine a baseline model (trained on all data) with fine-tuned models for specific particle types, selecting the best model per class based on validation metrics.

  • The improved system achieves a Wasserstein distance of 1.61 ± 0.02 (vs. 1.70–1.76 for baselines), with consistent improvements across all four metrics (MAEcond, dispersion ratio, Jaccard error), demonstrating that specialization outperforms a single generalist model.

6. Incorporate SHAP-based explainability for generative model validation

  • Use SHAP analysis on surrogate Bayesian neural networks (trained via knowledge distillation from the NF) to verify that generated outputs depend on physically meaningful input features (e.g., energy, momentum components) in expected ways.

  • The improved system can automatically audit whether a generative model has learned correct physical dependencies, catching implausible feature attributions before deployment, and can explain predictions for non-interpretable models like normalizing flows.

7. Enable fast inference via knowledge distillation to inverse autoregressive flows

  • Train an IAF student model to mimic the MAF teacher, achieving 420× faster sampling (0.38 ms vs. 160 ms per sample).

  • The improved system can be deployed in real-time or high-throughput simulation environments (e.g., online event filtering at LHC) while retaining the accuracy of the slower teacher model.

8. Handle class imbalance via dedicated fine-tuned models rather than data rebalancing

  • Instead of oversampling or reweighting, train separate fine-tuned models for underrepresented classes, which preserves the original data distribution and avoids introducing bias.

  • The improved system can generate accurate responses for rare classes (e.g., Σ+ with only 0.5% of data) that would otherwise be poorly modeled, without distorting the overall distribution.

9. Use bootstrap resampling for statistical significance testing of model improvements

  • Apply bootstrap at the unique-input level (preserving all associated responses) to compute p-values and confidence intervals for metric differences between models.

  • The improved system can determine whether performance gains are statistically robust or due to noise, enabling principled model selection in low-sample regimes.

10. Optimize for energy efficiency in model training

  • Choose fine-tuning over extended baseline training when performance is comparable, as it requires fewer total epochs and less compute.

  • The improved system can reduce carbon footprint and training costs by 50% while maintaining or improving accuracy, making it feasible for resource-constrained research environments.

Abstract

Simulating the ALICE Zero Degree Calorimeter (ZDC) neutron detector responses at the LHC is computationally expensive, requiring complex Monte Carlo chains. We develop a generative surrogate, focusing on Normalizing Flows (NFs). Through transfer learning, we pre-train on the full imbalanced dataset and fine-tune specialized models for different particle types (gamma, n,, K S 0,+) using two gradual-unfreezing schemes. As standard ZDC metrics like Wasserstein distance overlook conditional structure, we introduce refined metrics: conditional weighted MAE, dispersion ratio, and Jaccard co-activation error, that better capture physics-relevant input-output dependencies and response variability. Our ensemble of fine-tuned models achieves a Wasserstein distance of 1.61 plus or minus 0.02, outperforming baselines across all metrics. This work provides a generalizable NF-based framework for LHC detector simulation, combining NFs, conditional fine-tuning, and physics-motivated evaluation.

Sources

Related papers