SPORE: An Event-Level Sampling Pipeline for Multi-Telescope Neutrino Astronomy

arXiv:2608.11862 · astro-ph.IM, astro-ph.HE · Submitted 2026-08-12 · Read on arXiv

Jeffrey Lazar, Perrine Wilmet, Gwenhaël de Wasseige

Université catholique de Louvain

astro-ph.IM, astro-ph.HE

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 21 pages, 6 figures. Submitted to JCAP. Code available at https://github.com/jlazar17/spore

Code: https://github.com/jlazar17/spore

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: spore is an open-source Python package for simulating neutrino events from astrophysical point and extended sources using tabulated instrument response functions (IRFs).

Terminology

Summary

spore is an open-source Python package for simulating neutrino events from astrophysical point and extended sources using tabulated instrument response functions (IRFs). The package encodes three detector response components—effective area, point spread function, and energy resolution—in a selection-agnostic HDF5 format, which can represent a neutrino telescope whose response is supplied in that form, whether from a public release or a private study. Sampling algorithms cover point sources (inverse-CDF with Poisson or fixed-count modes), extended sky distributions (hierarchical inverse-CDF sampling, including full RA- and declination-dependent flux maps), and multi-detector joint analyses. The package is validated via a round-trip consistency test using the publicly available IceCube 10-year tracks data release: the released IRFs are ingested into the package and used to generate a synthetic event set, whose declination distribution reproduces the observed one to 10–15% across the northern sky. The reconstructed-energy distribution agrees to within about a third over the bulk of the sample but exceeds the data by up to a factor of three below 600 GeV, a discrepancy traced to the coarse true-energy binning of the public smearing matrix rather than to the sampling: an independent forward fold of the same IRFs reproduces it. The package is further compared against the IceCube HESE 7.5-year public data release: the sampled deposited-energy spectrum tracks the published best-fit expectation, and the observed data fall within the goodness-of-fit distribution built from 1,000 sampled pseudo-experiments, though on its well-fitting side, as expected for an expectation that was itself fit to those data.

Improvements for AI systems

Improvements to AI systems:

  1. Selection-agnostic detector response modeling: Implement a modular, HDF5-based encoding of instrument response functions (effective area, PSF, energy resolution) that allows AI systems to ingest and simulate detector data without hard-coded assumptions about the detector’s selection criteria. This enables rapid transfer of simulation pipelines across different neutrino telescopes or public data releases.

  2. Hierarchical inverse-CDF sampling for extended sources: Replace naive grid-based sampling with a hierarchical inverse-CDF algorithm that handles full RA- and declination-dependent flux maps. This improves computational efficiency and accuracy for AI systems that need to generate synthetic event sets from spatially non-uniform astrophysical sources.

  3. Multi-detector joint analysis framework: Add a sampling layer that supports simultaneous simulation across multiple detectors with distinct IRFs, enabling AI systems to perform combined likelihood analyses or train fusion models on multi-instrument data without re-engineering the sampling logic.

  4. Forward-folding validation for energy resolution: Incorporate an independent forward-fold check (comparing sampled reconstructed energies against a direct convolution of true-energy bins with the smearing matrix) to diagnose systematic biases in AI-generated synthetic data. This allows the AI system to automatically flag coarse binning or interpolation artifacts in public IRFs.

  5. Goodness-of-fit pseudo-experiment generation: Build a built-in routine that generates 1,000+ synthetic pseudo-experiments from the fitted model and computes a goodness-of-fit distribution. This enables AI systems to calibrate their own uncertainty estimates and detect overfitting (e.g., when observed data fall on the well-fitting side of the distribution, as in the HESE case).

  6. Declination-dependent synthetic data reproduction: Use the validated round-trip consistency test (10–15% agreement in declination distribution) as a benchmark metric for any AI system that generates or evaluates simulated neutrino data. This provides a quantitative target for improving generative models of detector outputs.

What the improved AI system can do:

  • Simulate realistic neutrino events from point or extended sources using any tabulated IRF, without requiring detector-specific code changes.

  • Automatically diagnose and correct systematic errors in public IRFs (e.g., coarse energy binning) by comparing forward-folded expectations to sampled outputs.

  • Generate calibrated pseudo-experiments for any fitted model, allowing robust hypothesis testing and uncertainty quantification in astrophysical searches.

  • Jointly analyze data from multiple neutrino telescopes (e.g., IceCube, KM3NeT) by sampling from a unified IRF format, enabling cross-instrument consistency checks and combined source searches.

  • Benchmark generative AI models against the declination and energy distributions from public data releases, with explicit tolerance thresholds (e.g., 10–15% in declination, factor-of-3 in low-energy reconstructed spectra) to flag model failures.

Abstract

We present an open-source Python package for simulating neutrino events from astrophysical point and extended sources using tabulated instrument response functions (IRFs). The package encodes three detector response components - effective area, point spread function, and energy resolution - in a selection-agnostic HDF5 format, which can represent a neutrino telescope whose response is supplied in that form, whether from a public release or a private study. Sampling algorithms cover point sources (inverse-CDF with Poisson or fixed-count modes), extended sky distributions (hierarchical inverse-CDF sampling, including full RA- and declination-dependent flux maps), and multi-detector joint analyses. We validate the framework via a round-trip consistency test using the publicly available IceCube 10-year tracks data release: the released IRFs are ingested into the package and used to generate a synthetic event set, whose declination distribution reproduces the observed one to 10-15% across the northern sky. The reconstructed-energy distribution agrees to within about a third over the bulk of the sample but exceeds the data by up to a factor of three below 600 GeV, a discrepancy we trace to the coarse true-energy binning of the public smearing matrix rather than to the sampling: an independent forward fold of the same IRFs reproduces it. We further compare against the IceCube HESE 7.5-year public data release: the sampled deposited-energy spectrum tracks the published best-fit expectation, and the observed data fall within the goodness-of-fit distribution built from 1,000 sampled pseudo-experiments, though on its well-fitting side, as expected for an expectation that was itself fit to those data.

Sources

Related papers