Pretrained Event Classification Model for High Energy Physics Analysis
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Pretrained Event Classification Model for High Energy Physics Analysis".
Jane: The paper was written by Joshua Ho, Ryan Roberts, Shuo Han and Haichen Wang from University of California, Berkeley and Lawrence Berkeley National Laboratory.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everyone! I'm Tom, and alongside me is the brilliant Jane. Today we're cracking open a fascinating new paper from the physics world, titled "Pretrained Event Classification Model for High Energy Physics Analysis."
Jane: And I'm Jane! Tom, I have to say, when I first saw this title, I got a little excited. It's basically taking the idea of foundation models—you know, like the big language models we hear about—and applying it to particle physics. Instead of training a new model for every single experiment, they're trying to build one general-purpose model that can be adapted.
Tom: Exactly! And that's a huge deal. For years, physicists have had to train machine learning models from scratch for every new analysis. That's expensive, it takes forever, and it doesn't always work well when you don't have enough data. This paper from the folks at UC Berkeley and Lawrence Berkeley National Lab is trying to change that whole game.
Jane: Right, and the authors—Joshua Ho, Ryan Roberts, Shuo Han, and Haichen Wang—they're not just theorizing about this. They actually built it. They trained a Graph Neural Network on one hundred twenty million simulated collision events covering twelve different physics processes. That's a massive dataset.
Tom: Massive is right. And the key insight here is that they're using a Graph Neural Network, which is perfect for this kind of data. Collision events are basically a bunch of particles flying out in different directions. You can think of each particle as a node in a graph, and the connections between them as edges. The network learns the relationships between those particles.
Jane: That's a great way to put it, Tom. It's like looking at a family photo and understanding the relationships—who's standing next to whom, who's in the background. The GNN picks up on those structural patterns, not just individual features. So when they fine-tune this model for a specific task, like distinguishing between two different types of Higgs boson production, it already has a head start.
Tom: And that head start is the whole point. We're going to get into the numbers and the results in a bit, but spoiler alert—the improvements are real, especially when you don't have a lot of training data. That's the scenario where this approach really shines.
Jane: I love that they're thinking about this as a foundation model, just like GPT or BERT in natural language. It's a paradigm shift for high-energy physics. Instead of reinventing the wheel for every analysis, you have this robust starting point that you can adapt.
Tom: Exactly. And the implications go beyond just saving time. It could make analyses more sensitive, which means we might catch subtle new physics signals that we'd otherwise miss. Stick around, because we're going to dig into the methods and the results next.
Jane: You won't want to miss it. Let's get into the details.
Summary: Tom: So we're back, and we're still talking about "Pretrained Event Classification Model for High Energy Physics Analysis." Jane, let's break down what this paper actually did, because the summary is pretty dense.
Jane: It really is, but the core idea is elegant. They built a foundation model for collider events. Think of a collider event as a snapshot of all the particles produced when protons smash together. The model takes all those particles—jets, electrons, muons, photons—and represents them as a graph. Each particle is a node with features like momentum and energy, and the edges capture the angular distances between them.
Tom: Right, and they pretrained this model on a huge dataset. We're talking about one hundred twenty million simulated events from twelve different physics processes. That includes all the major Higgs production mechanisms and a bunch of top quark processes. The goal of pretraining is to get the model to understand the general structure of collision events, not just one specific signature.
Jane: And they tried two different pretraining strategies. The first was multiclass classification—essentially, "given this event, which of the twelve processes produced it?" That forces the model to learn subtle differences between processes. The second was multilabel classification, where the model predicts things like the number of Higgs bosons or top quarks in the event, and their kinematic properties.
Tom: That's a clever setup. But here's the thing—after pretraining, they evaluated the model on seven different downstream tasks. Some of those tasks involved processes the model had never seen during pretraining, like supersymmetric particles or flavor-changing neutral currents. And they also tested it on ATLAS Open Data, which is real detector data from the ATLAS experiment at CERN.
Jane: That's the part that really impressed me. The ATLAS Open Data is processed through a completely different simulation chain than the Delphes fast simulation used in pretraining. So when the model performed well there, it showed that the learned representations are truly general, not just overfitted to one simulation setup.
Tom: And the results? Fine-tuning the pretrained model consistently beat training from scratch, especially in the low-data regime. We're talking about improvements of up to four percentage points in accuracy when you only have a few thousand training events. That's a big deal because in real physics analyses, you often don't have millions of events to work with.
Jane: Right, and the improvements shrink as you add more data, which makes sense. When you have a million events, you can learn from scratch pretty well. But when you're starved for data, having that pretrained foundation is a lifesaver.
Tom: So the summary is: pretraining works, it transfers across simulation frameworks, and it gives you a real edge when data is scarce. But there's a catch—the multilabel pretraining didn't work as well as the multiclass. We'll dig into why that is and what it means for the field next.
Jane: And we'll also look at how they figured out *why* the pretrained models work so well. It's not just about the numbers; they actually peeked inside the model's brain.
Improvements: Tom: Welcome back. We're still on "Pretrained Event Classification Model for High Energy Physics Analysis," and Jane, we just teased that the multiclass pretraining beat the multilabel approach. Let's talk about the improvements the paper suggests.
Jane: Right. So the paper is really pushing the idea that you don't need to train from scratch for every task. The improvements come in two flavors. First, there's the raw performance gain—fine-tuned models just do better, especially when you have limited data. Second, there's the computational efficiency gain—you reach a good solution much faster.
Tom: And that computational piece is huge. They measured something called "time-to-target," which is how long it takes to reach a certain accuracy threshold. At one hundred thousand training events, fine-tuning reached the target in just three to eight percent of the time it took the baseline model. That's a twelve-fold speedup or more.
Jane: That's staggering. But there's a nuance here. If you let the fine-tuned model train until it fully converges, it can actually take longer than the baseline at small sample sizes. That's because they use a lower learning rate for the pretrained layers to avoid destroying the learned features. So you trade a bit of wall-clock time for better final performance.
Tom: Right, but at full statistics—when you have a million or more events—the fine-tuned model trains faster overall. And the paper even calculates when the pretraining cost pays off. They spent about forty-five GPU hours on multiclass pretraining. If you're fine-tuning for a realistic analysis, you break even after about fourteen to fifty-two tasks, depending on your stopping criteria.
Jane: And that's not a crazy number. They point out that a real ATLAS measurement of Higgs couplings used forty-two separate classifiers. So the pretraining investment would pay for itself within a single large analysis. That's a compelling argument for the field to adopt this approach.
Tom: But it's not all sunshine. The multilabel pretraining actually hurt performance in some cases. For example, on the W H versus ZH task, multilabel pretraining made things worse at small sample sizes. The paper suggests that the multilabel objective was too prescriptive—it forced the model to focus on hand-designed labels like particle counts, which didn't align well with the downstream classification tasks.
Jane: That's a really important lesson. It's not enough to just pretrain on something physically motivated. The pretraining objective has to be flexible enough to learn general event structure. The multiclass approach, where the model has to discriminate between complete physics processes, naturally encourages that flexibility.
Tom: So the improvement isn't just "pretrain and win." It's "pretrain on the right task." And that's a subtle but crucial insight for anyone trying to build foundation models in this domain.
Jane: And we haven't even talked about the most interesting part—how they figured out *why* the fine-tuned models work. They used a technique called Centered Kernel Alignment to compare the internal representations of different models. That's coming up next, and it's a real eye-opener.
Tom: Can't wait. Let's keep going.
First Page: Tom: We're back, and we're diving into the first page of "Pretrained Event Classification Model for High Energy Physics Analysis." Jane, the introduction really sets the stage for why this work matters.
Jane: It does. The authors start by pointing out a fundamental problem in high-energy physics: every analysis trains its own machine learning model from scratch. That's inefficient, it requires specialized expertise, and it can lead to suboptimal performance when you don't have enough training data. They argue that foundation models—like GPT-four or BERT in other fields—offer a way out.
Tom: And they're not the first to think about foundation models for physics. There's been work on jet-level models, like OmniJet and MPM, that focus on individual jets. But this paper is different. It operates at the event level, looking at the whole collision event as a graph, not just individual jets.
Jane: Right, and they explicitly position their work as complementary to those jet-level models. They also mention Bumblebee, which is a Transformer-based event-level model, but it uses a generative pretraining objective. This paper uses a Graph Neural Network with discriminative pretraining—multiclass and multilabel classification—which is a fundamentally different approach.
Tom: And the scale is different too. They pretrained on one hundred twenty million events across twelve processes. That's substantially larger and more diverse than what most prior work has used. They're really trying to build a general-purpose model, not just something that works for one specific search.
Jane: The first page also sets up the evaluation strategy, which is really thorough. They test on seven tasks, including processes never seen during pretraining, and they include ATLAS Open Data to test cross-simulation generalizability. That's a rigorous benchmark.
Tom: And they tease the representational analysis with Centered Kernel Alignment. That's where they look inside the model to understand what's happening. We're going to get into that now, because it's the most fascinating part of the paper.
Jane: So the CKA analysis compares three things: the pretrained model before fine-tuning versus a baseline trained from scratch, the fine-tuned model versus the baseline, and the fine-tuned model versus the pretrained model. And the results are really revealing.
Tom: Let me guess—the encoders are similar, but the message-passing layers are totally different?
Jane: You nailed it. The encoder stages—where the model first embeds the node, edge, and global features—are nearly identical between the pretrained and baseline models. CKA scores around zero point nine to one point zero. That means pretraining already learns the same low-level feature extraction that a task-specific model would learn.
Tom: But then in the message-passing stages—where the model aggregates information across the graph—the similarity drops dramatically. We're talking CKA scores of zero point two to zero point five. That's like comparing two completely different ways of processing the graph structure.
Jane: And here's the kicker: after fine-tuning, the message-passing layers stay different from the baseline. The fine-tuned model doesn't converge to the baseline solution. It keeps its own unique pathway. But the final decoder—the part that produces the output—does align with the baseline after fine-tuning.
Tom: So the pretrained model finds a fundamentally different way to process the event graph, but it still ends up at the same decision boundary. And that's why it can beat the baseline—it's not just a better initialization, it's a different and richer computational strategy.
Jane: And the tasks where the decoder changes the most after fine-tuning are the same tasks where the performance gains are largest. That's a direct link between the representational analysis and the actual results. It's a beautiful story.
Tom: It really is. And it opens up a whole new way of thinking about how to design and evaluate foundation models for physics. Let's wrap this up in the conclusion.
Conclusion: Tom: Alright, we've reached the end of our discussion on "Pretrained Event Classification Model for High Energy Physics Analysis." Jane, let's put a bow on this.
Jane: Let's do it. This paper is a major step toward making foundation models a practical tool in high-energy physics. They built a Graph Neural Network, pretrained it on one hundred twenty million simulated events across twelve physics processes, and showed that fine-tuning it for specific tasks consistently beats training from scratch.
Tom: The gains are biggest when you have limited data—up to four percentage points in accuracy—and the model transfers across simulation frameworks, which they proved with ATLAS Open Data. That's not trivial.
Jane: And the computational story is compelling too. Fine-tuning reaches target performance in a fraction of the time, and the pretraining cost pays off after about fourteen to fifty-two tasks. For a field that trains dozens of classifiers per analysis, that's a game-changer.
Tom: But the most exciting part for me was the CKA analysis. They showed that pretraining doesn't just give you a better starting point—it gives you a fundamentally different way of processing the event graph. The fine-tuned model keeps its unique message-passing pathway while aligning its output with the task.
Jane: That's a deep insight. It suggests that pretraining on diverse physics processes teaches the model a general computational strategy, not just a set of weights. And that's why it generalizes so well.
Tom: And the lesson about pretraining objectives is important too. The multiclass approach worked, but the multilabel approach didn't. You have to choose the pretraining task carefully, or you can actually hurt downstream performance.
Jane: Exactly. This is a prototype, but it's a really strong one. The authors are clear that this is the first foundation model operating on collider final-state object data, and they've set a high bar for future work.
Tom: So what's next? Controlled studies on generator variations, better pretraining objectives, maybe even larger models. The possibilities are exciting.
Jane: Absolutely. This paper gives the community a solid foundation—pun intended—to build on. We'll be watching this space closely.
Tom: And that's a wrap on "Pretrained Event Classification Model for High Energy Physics Analysis." Thanks for joining us, and we'll see you on the next one.
Jane: Take care, everyone!
Joshua Ho, Ryan Roberts, Shuo Han, Haichen Wang
University of California, Berkeley · Lawrence Berkeley National Laboratory
hep-ph, cs.LG
Submitted: 2026-07-16
Comments: 13 pages, 2 figures
Journal ref: JINST 21 (2026) P08006
DOI: 10.1088/1748-0221/21/08/P08006
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 70/100
The gist: This paper introduces a foundation model for event classification in high-energy physics, built on a Graph Neural Network (GNN) architecture and trained on 120 million simulated proton-proton
Key concepts
- Graph Neural Network (GNN)
- A GNN is used because collision events are represented as graphs where each particle is a node and connections between them are edges. This structure allows the network to learn relationships between particles, capturing structural patterns in the event data that individual features alone might miss.
- Pretrained Model
- This involves training a general-purpose model on a massive dataset of simulated collision events across twelve different physics processes. The goal is to create a robust starting point that can be adapted to new, specific analysis tasks without retraining from scratch for every experiment.
- Centered Kernel Alignment (CKA)
- This technique is used to compare the internal representations of different models. The analysis showed that while the initial feature extraction layers are similar between pretrained and baseline models, the message-passing layers—which process graph information—are fundamentally different, showing a richer computational strategy.
- Time-to-Target
- This metric measures how long it takes for a model to reach a specific accuracy threshold. Fine-tuning the pretrained model can reach this target much faster than the baseline model, demonstrating significant computational efficiency gains in reaching required performance levels.
Terminology
Summary
This paper introduces a foundation model for event classification in high-energy physics, built on a Graph Neural Network (GNN) architecture and trained on 120 million simulated proton-proton collision events spanning 12 distinct physics processes. The model is pretrained to learn a general and robust representation of collision data using challenging multiclass and multilabel classification tasks.
The pretraining dataset consists of approximately 120 million events, evenly distributed across 12 distinct physics processes, including all major Higgs boson production mechanisms (gluon fusion production, vector boson fusion, associated production of the Higgs boson with a W boson or a Z boson, associated production of the Higgs boson with a top-quark pair, and associated production of the Higgs boson with a single top quark and a forward quark) and six top quark production processes (single top production, top-quark pair production, top quark pair production in association with a pair of photons, associated production of a top-quark pair with a W boson, simultaneous production of three top quarks, and simultaneous production of four top quarks). The events are generated using Madgraph@NLO 2.7.3 at next-to-leading order in Quantum Chromodynamics, processed through Pythia 8.235 for parton showering and heavy particle decays, followed by Delphes 3.4.2 configured to emulate the ATLAS detector for fast detector simulation. The center of mass energy of the proton-proton collision is set to 13 TeV.
The GNN architecture naturally accommodates the point-cloud structure of particle physics data, employing the DGL framework with a PyTorch backend. A fully connected graph is constructed for each event, with nodes corresponding to reconstructed jets, electrons, muons, photons, and missing transverse energy. The features of each node include the four-momentum (pT, η, ϕ, E) of the object with a massless assumption, the b-tagging label (for jets), the charge (for leptons), and an integer labeling the type of object represented by the node. The angular distances (Δη, Δϕ, ΔR) are assigned as edge features and the number of nodes N in the graph as a global feature. The model is comprised of three primary components: an encoder, the graph network, and a decoder. In the encoder, three MLPs embed the nodes, edges, and global features into a latent space of dimension 64. The graph network block performs an edge update, followed by a node update, and finally a global update, and this graph block is iterated four times with the same update MLPs. Each MLP consists of 4 linear layers, each with an output width of 64, with the ReLU activation function, and the output of the MLP is passed through a LayerNorm layer. The total number of trainable parameters in this model is about 400,000.
Two complementary pretraining approaches are explored: multiclass classification, which trains the model to distinguish between 12 different physics processes simultaneously using categorical cross entropy as the loss function, and multilabel classification, which combines both classification and regression tasks to characterize collision events using a comprehensive set of 41 labels that capture both particle multiplicities and kinematic properties. The multilabel approach uses binary cross-entropy for classification labels and mean-squared error for regression labels, with the final loss computed as an equally-weighted average across all labels.
The model's performance is evaluated across seven event classification tasks, which include new physics processes not encountered during pretraining as well as ATLAS Open Data to demonstrate generalizability across different simulation frameworks, from Delphes fast simulation to full ATLAS detector simulation. The seven tasks are: tt̄H(→ γγ) with CP-even versus CP-odd t-H interaction (Delphes), tt̄ with FCNC top quark decays versus tHq processes (Delphes), tt̄W versus tt̄t processes (Delphes), s-top pair production with Higgs bosons in the decay chain versus tt̄H processes (Delphes), W H versus ZH production modes (Delphes), 5-class multiclass classification of Higgs production modes where the Higgs decays exclusively to two photons (ATLAS Open Data), and 3-class multiclass classification of triboson events (ATLAS Open Data).
Fine-tuning the pretrained model significantly improves classification performance, particularly in scenarios with limited training data, demonstrating gains in both accuracy and computational efficiency. The paper states: "In general, the fine-tuned pretrained model achieves at least the same level of classification performance as the baseline model. Notably, there are significant improvements, particularly when the sample size is small, ranging from 103 to 105 events. In some cases, the accuracy (AUC) improvements exceed 4 percentage points (2.5 AUC points), demonstrating that pretrained models provide a strong initial representation that compensates for limited data." As the training sample size grows to 106 and eventually 107 events, the added benefit of pretraining diminishes, with models trained from scratch approaching or even matching the accuracy of fine-tuned pretrained models. The paper notes that training sample sizes of 103 to 105 events per class reflect realistic operating conditions in HEP analyses.
Regarding the two pretraining approaches, the paper states: "Although both pretraining approaches offer benefits, multiclass pretraining tends to provide more consistent improvements across tasks, especially in the low-data regime. In contrast, multilabel pretraining can sometimes lead to neutral or even slightly negative effects for certain tasks and data sizes. The paper interprets this as evidence that
not all physically motivated supervised pretraining objectives produce equally transferable representations and that
overly prescriptive auxiliary labels can bias the learned representation toward features that are not optimally aligned with downstream event-classification tasks."
For the ATLAS Open Data tasks, the paper reports: "For the Higgs production task, multiclass pretraining yields no improvement with −0.1 ± 0.2 percentage points in accuracy and a small increase of +0.1 ± 0.0 AUC points over the baseline, while multilabel pretraining leads to a slight degradation in both metrics. For the triboson task, multiclass pretraining provides a more substantial improvement of +3.7 ± 2.3 percentage points in accuracy and +4.6 ± 3.3 AUC points. The paper concludes that
the consistent improvement from multiclass pretraining across both simulation frameworks suggests that the pretrained model learns physics-driven representations that are robust to differences in detector simulation."
To investigate the underlying mechanisms behind these performance improvements, the paper employs a representational similarity evaluation framework based on Centered Kernel Alignment (CKA). The layer-wise CKA analysis reveals three distinct regimes. In the encoding stages (Node Encoder, Edge Encoder, Global Encoder), CKA values are consistently high across all tasks, in the range of 0.9–1.0, indicating that the pretrained model has already developed low-level feature representations that closely resemble those of the best baseline. In the message passing stages (Node Update, Edge Update, Global Update), the similarity drops dramatically to values in the range of 0.2–0.5 across all tasks, indicating that the pretrained and baseline models process and aggregate graph information through fundamentally different computational pathways. At the Global Decoder, the similarity is task-dependent and correlates with task difficulty, with harder tasks showing substantially different representations from the baseline. The paper summarizes: Fine-tuning preserves general-purpose encoders while developing fundamentally different message-passing pathways to arrive at superior task performance.
The three comparisons together reveal that "the pretrained model begins with encoder representations already closely aligned with the best baseline, which is a strong initialization advantage. It then processes graph information through a fundamentally different message passing pathway that is retained through fine-tuning... Fine-tuning then primarily reorganizes the final decoder representations to align with the downstream task, leaving the encoder and message passing stages largely unchanged."
The computational efficiency analysis uses two metrics: time-to-target (wall-clock time required to reach a test AUC within 10−3 of the baseline model's best test AUC) and full training time (time required to reach a stopping condition of 5 consecutive epochs with no improvement in training loss). The paper reports: At 105 events, fine-tuning reaches the AUC target in just 3–8% of the baseline training time across all tasks, corresponding to speedups exceeding 12× for every task.
Under the full-training-time metric, at full statistics the full-training-time ratio falls below unity for every task, with a 7-task mean of 64.6%.
The pretraining costs are 45.5 GPU hours for multiclass pretraining and 60.0 GPU hours for multilabel pretraining. The paper computes that the pretraining cost recovered after approximately 14–52 finetuned tasks depending on the stopping condition applied,
and notes that the ATLAS measurement of Higgs boson couplings using the H → γγ decay channel involved training 42 classifiers for event categorization,
suggesting the foundation model approach can meaningfully reduce total computational cost of a realistic high-energy physics analysis.
The paper also reports on hyperparameter optimization: "A total of 400 GPU hours on a single NVIDIA A100 were invested for systematic hyperparameter optimization... These results show that the GNN architecture exhibits stability across hyperparameter configurations: excluding a single outlier, variations in the test AUC remain within ±0.2%."
Improvements for AI systems
Based on the paper, I can implement the following specific improvements to an AI system for high-energy physics event classification:
Improvement: Replace task-specific models trained from scratch with a Graph Neural Network pretrained on 120 million simulated events across 12 physics processes using multiclass classification.
Capability: The system can achieve:
-
+2.9 to +4.6 percentage points accuracy improvement on binary classification tasks with only 103–105 training samples
-
+3.7 percentage points accuracy improvement on ATLAS Open Data triboson classification with only 300,000 samples
-
Consistent performance gains across 7 downstream tasks, including processes never seen during pretraining
Sources
- GPT-4 Technical Report
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- High-Resolution Image Synthesis with Latent Diffusion Models
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- How transferable are features in deep neural networks?
- Bumblebee: Foundation Model for Particle Physics Discovery
- Learning Symmetry-Independent Jet Representations via Jet-Based Joint Embedding Predictive Architecture
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Solving Key Challenges in Collider Physics with Foundation Models
- Re-Simulation-based Self-Supervised Learning for Pre-Training Foundation Models
- Finetuning Foundation Models for Joint Analysis Optimization
- Point cloud-based diffusion models for the Electron-Ion Collider
- A Language Model for Particle Tracking
- Xiwu: A Basis Flexible and Learnable LLM for High Energy Physics
- Is Tokenization Needed for Masked Particle Modelling?
- Observation of four-top-quark production in the multilepton final state with the ATLAS detector
- The automated computation of tree-level and next-to-leading order differential cross sections, and their matching to parton shower simulations
- A framework for Higgs characterisation
- Automatic spin-entangled decays of heavy resonances in Monte Carlo simulations
- An Introduction to PYTHIA 8.2
Related papers
- Classification of g-modes for neutron stars with a strong transition: Novel universal relation including slow stable hybrid stars
- Higgsino Dark Matter Interpretation of the LUX-ZEPLIN 248 keV Nuclear-Recoil Event
- A Unified Bogoliubov Approach to Primordial Gravitational Waves: From Inflation to Reheating
- Probing Memory-Burdened Primordial Black Holes with High-Energy Neutrinos
- Enhanced Dark Matter Quantum Sensing via Phase-Space Geometric Interferometry
- Axions as Dark Matter, Dark Energy, and Dark Radiation