Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
hep-ph, cs.LG, hep-ex, physics.data-an
Submitted: 2024-01-24
Updated: 2024-07-11
Journal ref: Machine Learning: Science and Technology 5, 035074 (2024)
Code: https://github.com/rodem-hep/mpm
License: http://creativecommons.org/licenses/by/4.0/
The gist: We propose masked particle modeling (MPM) as a self-supervised method for learning generic, transferable, and reusable representations on unordered sets of inputs for use in high energy physics (HEP)
Terminology
Abstract
We propose masked particle modeling (MPM) as a self-supervised method for learning generic, transferable, and reusable representations on unordered sets of inputs for use in high energy physics (HEP) scientific data. This work provides a novel scheme to perform masked modeling based pre-training to learn permutation invariant functions on sets. More generally, this work provides a step towards building large foundation models for HEP that can be generically pre-trained with self-supervised learning and later fine-tuned for a variety of down-stream tasks. In MPM, particles in a set are masked and the training objective is to recover their identity, as defined by a discretized token representation of a pre-trained vector quantized variational autoencoder. We study the efficacy of the method in samples of high energy jets at collider physics experiments, including studies on the impact of discretization, permutation invariance, and ordering. We also study the fine-tuning capability of the model, showing that it can be adapted to tasks such as supervised and weakly supervised jet classification, and that the model can transfer efficiently with small fine-tuning data sets to new classes and new data domains.
Sources
- On the Opportunities and Risks of Foundation Models
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- GPT-4 Technical Report
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- BEiT: BERT Pre-Training of Image Transformers
- Zero-Shot Text-to-Image Generation
- AstroCLIP: A Cross-Modal Foundation Model for Galaxies
- Towards Galaxy Foundation Models with Hybrid Contrastive Learning
- Anomalies, Representations, and Self-Supervision
- Pre-training strategy using real particle collision data for event classification in collider physics
- Particle Transformer for Jet Tagging
- Point Cloud Generation using Transformer Encoders and Normalising Flows
- Differentiable Vertex Fitting for Jet Flavour Tagging
- Equivariant Transformer is all you need
- Attention to Mean-Fields for Particle Cloud Generation
- $\nu^2$-Flows: Fast and improved neutrino reconstruction in multi-neutrino final states with conditional normalizing flows
- Jet Diffusion versus JetGPT -- Modern Networks for the LHC
- Finetuning Foundation Models for Joint Analysis Optimization
- Neural Discrete Representation Learning
Related papers
- Classification of g-modes for neutron stars with a strong transition: Novel universal relation including slow stable hybrid stars
- Higgsino Dark Matter Interpretation of the LUX-ZEPLIN 248 keV Nuclear-Recoil Event
- A Unified Bogoliubov Approach to Primordial Gravitational Waves: From Inflation to Reheating
- Probing Memory-Burdened Primordial Black Holes with High-Energy Neutrinos
- Enhanced Dark Matter Quantum Sensing via Phase-Space Geometric Interferometry
- Axions as Dark Matter, Dark Energy, and Dark Radiation