Towards foundation-style models for energy-frontier heterogeneous neutrino detectors via self-supervised pre-training

arXiv:2604.07037 · hep-ex, cs.CV · Submitted 2026-04-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Towards foundation-style models for energy-frontier heterogeneous neutrino detectors via self-supervised pre-training".

Jane: The gist: Self-supervised pre-training, combining masked autoencoder reconstruction with relational voxel-level objectives, enables a sparse Vision Transformer framework to learn reusable representations from heterogeneous detector data for energy-frontier neutrino detectors.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Alright, so what's the main thrust of "Towards foundation-style models for energy-frontier heterogeneous neutrino detectors via self-supervised pre-training"? The paper argues that traditional event interpretation is impractical at the TeV scale because the signatures are just too dense and overlapping.

Jane: They propose a sparse Vision Transformer framework to solve this by learning reusable representations directly from heterogeneous detector data, which is tough when labeled data is scarce or analyses need to cover many different physics goals.

Lu: The thesis involves a two-phase pre-training strategy: first, masked autoencoder reconstruction alone, and then adding relational voxel-level targets for things like ghost identification and particle categorization.

Meng: They introduce a sparse encoder design that mixes sparse convolutional patch embeddings with module-aware self-attention and Perceiver-IO fusion to handle the different detector streams efficiently.

Lalam: The pre-training uses masked reconstruction to map decoder tokens back to voxels, and then adds a relational forward pass where the encoder predicts hierarchy labels, ghost labels, and particle categories on kept three deeCal patches <ref:2604.07037#pg2>.

Tom: So why does this matter for the physics side? They claim that this composite objective improves downstream performance beyond just using masked reconstruction alone, with the largest gains in the most challenging channels.

Jane: And they show that this learned representation is useful across a suite of tasks, including flavour identification, charm-quark identification, and even event kinematics and vertex reconstruction after fine-tuning.

Lu: What's interesting is that they demonstrate this representation transfers beyond the source domain to public benchmarks covering different detector technologies and energy regimes.

Meng: That transfer capability suggests the learned structure isn't just tailored to one specific experiment but captures something more general about dense interaction topologies.

Conclusion: Tom: So, looking at "Towards foundation-style models for energy-frontier heterogeneous neutrino detectors via self-supervised pre-training," the title itself points to building foundational models specifically for these complex, high-energy neutrino detectors using self-supervised training.

Jane: The authors are Sa´ul Alonso-Monsalve, Fabio Cufino, Umut Kose, Anna Mascellani, and Andr´e Rubbia from ETH Zurich. They focus on creating a sparse Vision Transformer framework that learns reusable representations from heterogeneous detector data.

Lu: What this means simply is that they've laid out a path toward general detector encoders for neutrino physics by using masked autoencoder reconstruction alongside relational voxel-level objectives to learn structure at the fundamental level.

Meng: For practical applications, it suggests we can use these pre-trained encoders to start with initial classifications or reconstructions on new detector data much faster than training from scratch.

Lalam: It sets up a concrete path toward more general detector encoders and motivates further work on hybrid pre-training objectives and domain adaptation for these complex environments.

IPA, ETH Zurich

hep-ex, cs.CV

Submitted: 2026-04-08

Updated: 2026-08-12

Comments: 18 pages, 6 figures

Journal ref: Nature Machine Intelligence (2026)

DOI: 10.1038/s42256-026-01309-6

Code: https://github.com/rubbiaa/FASER

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: The gist: Self-supervised pre-training, combining masked autoencoder reconstruction with relational voxel-level objectives, enables a sparse Vision Transformer framework to learn reusable

Key concepts

Sparse Vision Transformer (ViT)
A specific neural network architecture designed to process complex, high-dimensional data efficiently. This framework uses a sparse approach combined with self-attention mechanisms to handle the diverse inputs from different detector components like calorimeters and trackers without requiring dense connections across all data points.
Masked Autoencoder (MAE)
A self-supervised pre-training technique where the model learns to reconstruct missing parts of an input. In this work, it is used to train the encoder by masking patches of detector data and forcing the model to predict what those missing parts should look like, establishing a strong initial representation.
Relational Voxel-level Objectives
Additional training targets that go beyond simple reconstruction. These objectives teach the model relationships between different parts of the detector, such as predicting voxel labels or particle categories based on surrounding information. This helps the model understand the physical structure and context of neutrino interactions.
Heterogeneous Detector Data Fusion
The process of combining data from different detector types—like 3DCal, AHCAL, ECAL, and muon spectrometer—into a single representation. The framework uses techniques like Perceiver-IO fusion to integrate these distinct data streams into a unified set of tokens for the Transformer encoder.

Terminology

Summary

The gist: Self-supervised pre-training, combining masked autoencoder reconstruction with relational voxel-level objectives, enables a sparse Vision Transformer framework to learn reusable representations from heterogeneous detector data for energy-frontier neutrino detectors.

Introduction and Motivation

Accelerator-based neutrino physics is entering an energy frontier regime in which interactions reach the TeV scale and produce exceptionally dense, overlapping detector signatures Event interpretation is impractical for conventional reconstruction and challenging for supervised models trained from scratch, particularly when labelled data are scarce and analyses span diverse objectives The challenge is not simply whether learned models outperform existing pipelines, but whether any practical analysis of these events is feasible without them This is precisely the setting in which representation learning becomes compelling

The Framework and Architecture

The proposed framework utilizes a sparse Vision Transformer (ViT)-like architecture that combines masked autoencoder (MAE) pre-training with relational voxel-level objectives available from simulation The encoder is later fine-tuned on flavour identification, charmed-quark identification, event kinematics and vertex reconstruction The framework introduces a sparse encoder for heterogeneous detector data that combines sparse convolutional patch embeddings, module-aware self-attention, and Perceiver-IO fusion across calorimetric and tracking streams

The input representation involves converting native detector representations into detector-specific tokens The 3DCal is tokenised in patches of 12 × 12 × 10 voxels, yielding a 4×4×20 patch grid The AHCAL is tokenised in patches of 6 × 6 × 5 voxels (3 × 3 × 8 grid, up to 72 tokens) The ECAL is encoded as a compact token derived from its energy matrix The muon spectrometer input is variable length: track or hit-plane summaries are embedded and pooled into a compact spectrometer token A hierarchical encoder performs module-level self-attention before Perceiver-IO fusion of the heterogeneous detector streams

Self-Supervised Pre-training Strategy

The pre-training strategy is a two-phase schedule combining masked reconstruction with relational voxel-level targets In the first phase, the model is trained for 400 epochs using masked reconstruction alone To map a single decoder token back to the many voxels within its patch, a multi-rank separable basis (initialised from discrete cosine transform (DCT) coefficients) is used In the second phase, training continues from this checkpoint for 100 additional epochs while introducing, with probability 0.5 per batch, a relational forward pass in which the encoder predicts voxel-level ghost labels, hierarchy labels and particle-category labels on kept 3DCal patches (with a lower mask ratio of 0.25) The ghost target is binary, but the two semantic targets are not hard one-hot labels: each reconstructed voxel inherits all matched truth contributions, weighted by their fractional contribution to that voxel and normalised across classes

Downstream Fine-Tuning and Evaluation

The pre-trained encoder is then fine-tuned jointly on event-level classification and regression tasks This joint fine-tuning tests whether the pre-trained representation is useful as a shared basis for multiple measurements rather than for a single optimised classifier Downstream targets include neutrino flavour, charmedquark category, visible momentum, jet momentum, and the primary interaction vertex Performance is evaluated with one-vs-rest ROC curves and a figure of merit FOM = S/√S + B

Key Findings

The results demonstrate that self-supervised pre-training becomes most valuable precisely where energy-frontier neutrino events are hardest to interpret The MAE+Rel objective provides the largest gains for the less abundant and more topologically complex signatures For the dominant channels, the gains are consistent, with MAE+Rel showing improvements in areas like ντ CC → had where jet reconstruction is especially challenging The relational objectives disproportionately improve the lower-yield channels that are most critical for the detector’s physics reach

The learned representation improves performance and data efficiency across a multi-task fine-tuning suite, and transfers beyond the source domain to publicly available benchmarks spanning different detector technologies and energy regimes The transfer results place the work in a broader context by showing that the pre-training captures structure that is useful well beyond the original detector, task and energy range

The study concludes that for dense neutrino events, reconstruction-style objectives are a strong starting point but not the whole answer The results show several ingredients usually discussed separately can coexist in one detector-aware sparse encoder The present work does not claim that a general detector foundation model has been achieved

Conclusion

The study establishes the representation-learning basis on which such studies could be built for energy-frontier neutrino physics It outlines a concrete path towards more general detector encoders and motivates further work on hybrid pre-training objectives and domain adaptation

--- Page 22 ---

TABLE 2: Model size and active computational cost. The downstream count refers to the model used for fine-tuning and inference in the headline FASERCal experiments. The pre-training count includes the masked-autoencoder decoder and relational heads, which are discarded before downstream fine-tuning

--- Page 20 ---

FIG. 3: Event display of a simulated charged-current muon antineutrino interaction in the FASERCal detector. The event is a high-energy ¯νµ CC interaction with a true visible energy of Evis ≃ 1265 GeV and a reconstructed primary-lepton momentum of pµ ≃ 745 GeV

--- Page 21 ---

FIG. 2: Supplementary class-resolved regression boxplots. Regression-error boxplots for the selected event sample, resolved by true flavour category and target observable

--- Page 10 ---

FIG. 4: Data-efficiency study for the three initialisation strategies. Performance of Scratch, MAE and MAE+Rel across labelled training budgets from 100 to 100,000 events

--- Page 8 ---

TABLE I: Generator robustness and cross-domain transfer benchmarks. A single table collects the three external-validation summaries

--- Page 13 ---

TABLE 1: Training hyperparameters. Pre-training proceeds in two phases: Phase 1 trains with masked reconstruction only (MAE), and Phase 2 continues from the Phase 1 checkpoint with an additional relational pass (MAE+Rel)

--- Page 5 ---

FIG. 1: Example events from the test set illustrating masked reconstruction and relational tasks in the 3DCal detector. For each event, the top panel displays the ground-truth detector readout

--- Page 6 ---

FIG. 3: Attribution and representation diagnostics for the learned representation. A Two example patch-level 3DCal saliency maps for test charged-current events from the fine-tuned MAE+Rel

--- Page 20 ---

FIG. 5: Overview of the framework. A FASERCal case study, with the 3DCal as the primary interaction volume followed by the ECAL, AHCAL and muon spectrometer

--- Page 14 ---

Appendix B: Supplementary architectural details

  1. Tokenisation and latent structure

In the configuration used for the main experiments, the 3DCal is tokenised into 12 × 12 × 10 voxel patches, yielding a 4×4×20 patch grid

--- Page 19 ---

Appendix D: Supplementary target and loss definitions

  1. Pre-training targets

Masked reconstruction predicts voxel occupancy and charge for masked 3DCal and AHCAL patches, together with masked ECAL and muon-spectrometer summaries when those inputs are dropped

--- Page 20 ---

Appendix C: Supplementary training details

Improvements for AI systems

  1. Improve event interpretation in high-energy neutrino physics by employing a sparse Vision Transformer (ViT)-like framework that combines masked autoencoder (MAE) pre-training with relational voxel-level objectives. This system can perform flavour identification, charmed-quark identification, momentum regression and vertex reconstruction on simulated FASERCal events, achieving performance comparable to scratch training with an order of magnitude more data when only a few hundred to a few thousand labelled events are available.

  2. Develop reusable detector representations by utilizing a multimodal pre-training strategy that augments masked reconstruction with relational voxel-level targets (ghost identification, interaction hierarchy and particle-category labels). This allows the system to learn meaningful spatial and semantic correlations before any downstream fine-tuning is applied, which is particularly effective in channels where dense overlap, secondary activity and partial containment complicate classification.

  3. Enhance data efficiency by leveraging the pre-trained encoder, as it demonstrates that MAE+Rel gives the strongest overall performance across the study even at small label budgets. This means a system can achieve flavour-classification performance exceeding Scratch trained on roughly 104 events with only about 103 labelled events, making rare event studies more practical.

  4. Enable cross-domain transfer by utilizing the learned representation to perform well on public datasets outside the source domain, such as plastic scintillator benchmarks and PILArNet (LArTPC). This allows the system to adapt its knowledge from one detector technology or energy scale to another, suggesting that the pre-training captures structure that is useful well beyond the original detector.

  5. Increase robustness against systematic uncertainties by implementing a coherent global energy-scale stress test on the learned representation, which shows no acute brittleness to an O(10%) coherent calorimeter-scale bias. This means the system can maintain stable performance even when faced with detector miscalibration or scale shifts during inference.

  6. Establish physically motivated latent spaces by using attribution and representation diagnostics, which show that 3DCal provides the backbone of the event interpretation while auxiliary branches contribute in physically plausible, channel-dependent ways. This allows for a diagnostic understanding of model reliance, identifying which detector inputs are used for specific physics tasks.

  7. Create flexible deployment strategies by using detector-subsystem ablations to understand how information is utilized, confirming that suppressing a branch is not the same as modelling a biased or inefficient branch. This informs the design of future systems by suggesting that a realistic deployment programme should include subsystem-specific perturbations to test robustness against real-world inefficiencies.

Abstract

Accelerator-based neutrino physics is entering an energy-frontier regime in which interactions reach the TeV scale and produce exceptionally dense, overlapping detector signatures. In this regime, event interpretation becomes impractical for conventional reconstruction approaches, particularly when labelled data are scarce and the analysis spans diverse downstream objectives. We present a sparse ViT framework for learning reusable representations from heterogeneous detector data. Self-supervised pre-training combines masked autoencoder reconstruction with relational voxel-level objectives for hierarchy, ghost and particle identification, and the resulting shared encoder is then jointly fine-tuned across classification and regression tasks. Evaluated on simulated events from the proposed FASERCal concept at the LHC, we find that pre-training consistently improves neutrino flavour and charm-quark identification, momentum regression, and vertex reconstruction over training from scratch, with the addition of relational objectives yielding further gains in the most topologically complex channels. Interpretability analyses further show that pre-training yields a more structured latent space, while detector-subsystem ablations recover physically plausible channel-dependent roles for the heterogeneous inputs. A data-efficiency study shows that, with roughly 10 cubed labelled events, the pre-trained encoder already matches the flavour-classification performance of a randomly initialised model trained on an order of magnitude more data. The learned representations also transfer effectively to publicly available benchmarks spanning different detector technologies and energy scales, matching or exceeding published baselines. These results support self-supervised pre-training on multimodal detector data as a scalable route towards reusable representations for neutrino and particle-detector analysis.

Sources

Related papers