Generating particle physics Lagrangians with transformers

arXiv:2501.09729 · cs.LG, cs.SC, hep-ph, hep-th · Submitted 2025-01-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Generating particle physics Lagrangians with transformers".

Jane: The paper was written by Yong Sheng Koay, Rikard Enberg, Stefano Moretti and Eliel Camargo-Molina from Uppsala University and University of Southampton.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone! Today we're looking at a paper that's got me genuinely pumped: "Generating particle physics Lagrangians with transformers." Jane, this title is a mouthful, but it's basically about teaching an AI to write the equations that describe the universe.

Jane: And I love it because it sounds terrifyingly complex, but the core idea is so elegant. A Lagrangian, in simple terms, is like the instruction manual for a physical system. It tells you how particles move and interact. Physicists spend years learning how to build these things correctly.

Tom: Right, and normally, you'd need a human expert to look at a list of particles and figure out all the allowed interactions. This team from Uppsala University trained a transformer model—the same kind of architecture that powers ChatGPT—to do that automatically. They fed it a list of particles, and it spits out the Lagrangian.

Jane: And it's not just memorizing. They show it can handle the Standard Model's specific symmetries, which are these mathematical rules about how particles transform. It’s like teaching a model the grammar of the universe, not just a bunch of sentences.

Lu: The exciting part for me is that this isn't just a trick. The paper shows the model has internalized concepts like group representations. It's not just matching patterns; it's learning the abstract structure of the physics. That’s a huge step for symbolic AI.

Meng: But from an engineering standpoint, I have to ask: why not just use the code that generated the training data? If you have a pipeline that can write these Lagrangians automatically, why train a model to do it?

Tom: That’s the million-dollar question, Meng. The authors are upfront about it. The goal isn't to replace the code. It's to build a foundation model that can eventually do more—like explore theories beyond the Standard Model or even connect experimental data to theoretical models.

Jane: So it's a first step toward a much bigger vision. And the fact that they’re making the model and datasets public means other researchers can build on this. That's how science moves forward.

Lu: Exactly. And the performance is remarkable. Over ninety percent accuracy on test data. It’s not perfect, but it’s a proof of concept that transformers can handle this kind of symbolic, symmetry-rich mathematics.

Tom: And that's the hook for our next segment—we're going to dig into how they actually built this thing and what the results really mean. Stay with us!

Summary: Tom: We're back with "Generating particle physics Lagrangians with transformers." Jane, we teased the big picture, but let's get into the nitty-gritty of what the paper actually reports.

Jane: So they built a pipeline using a tool called AutoEFT to generate a massive dataset of Lagrangians. Then they trained a BART model—that's a specific type of transformer—with about three hundred fifty-seven million parameters. The model's job is to take a list of fields and produce the correct Lagrangian.

Tom: And the results are solid. On their test datasets, they’re getting perfect Lagrangian scores over ninety percent of the time. That means the model is writing the correct terms with the correct contractions, and not adding extra junk.

Meng: I was looking at the numbers. The "sampled" model, trained on a dataset skewed toward simpler examples, hit ninety-two percent on the merged test set. The "uniform" model, trained on a more balanced set, hit ninety-three point two percent. But they fail differently.

Lu: Right, the sampled model is better at getting the terms right, but it sometimes adds extra ones. The uniform model is more conservative. It’s a classic trade-off between precision and recall, but in a symbolic domain.

Jane: And they didn't just test on random data. They tested on real physics models, like the Standard Model itself. The model gets most of it right, but it struggles with the full Standard Model because it has six Yukawa interactions, and the model tends to miss a few.

Tom: That’s a great point. It’s not a perfect tool yet. But the fact that it can handle the lepton sector or the quark sector almost flawlessly is impressive. It’s learning the rules, not just memorizing examples.

Meng: The other thing that stood out to me is the out-of-distribution testing. They pushed the model to handle up to ten fields, even though it only saw six during training. It still produces reasonable Lagrangians, though accuracy drops.

Lu: That’s the key evidence that it’s generalizing. If it were just memorizing, it would fall apart completely on unseen numbers of fields. Instead, it degrades gracefully. It’s missing terms, but it’s not writing gibberish.

Jane: And that graceful degradation is what makes me excited about the future. It shows the model has a conceptual understanding, not just pattern matching. Next, we’re going to look at the specific improvements they suggest to make it even better.

Tom: And that’s where the real potential lies. Stick around.

Improvements: Tom: Welcome back. We're still on "Generating particle physics Lagrangians with transformers." Jane, we've seen the model works, but the paper is honest about its flaws. What do they suggest we do about it?

Jane: The main issue is counting. When you give the model more than six or seven fields, it starts to lose track of how many it has. It’s a known weakness of the bidirectional encoder in BART. It can’t count tokens in a sequence as well as a causal decoder.

Meng: So the fix is architectural. They suggest either a different architecture or a specialized counting mechanism. That’s a concrete engineering challenge.

Lu: But there's also a data solution. They noticed the model struggles with Yukawa terms—those are interactions between two fermions and a scalar. They’re rare in the dataset, so the model doesn't see enough examples. Enriching the training data with more of those would help.

Tom: And they also mention the tokenization of numbers. U(one) hypercharges are fractions, and the model has trouble with out-of-distribution fractions, like two/four instead of one/two. That’s a representation problem.

Jane: Right. They suggest a continuous number encoding instead of text-based tokens. That could help the model understand arithmetic, not just pattern-match.

Meng: So we have three levers: architecture, data distribution, and tokenization. Which one do you think gives the biggest bang for the buck?

Lu: I’d say the tokenization. If the model can't understand numbers, it can't conserve charge. That’s fundamental. The architecture can be tweaked, but if the input is fundamentally ambiguous, you're stuck.

Jane: And they’re already thinking about scaling up. They want to add more complex symmetries, flavor generations, discrete symmetries. All of that is possible with their tokenization scheme, but it needs the model to be more robust.

Tom: So the improvements aren't just about fixing bugs; they're about building a foundation for a much bigger model. That leads us to the first page of the paper, where they lay out their grand vision.

Jane: And that vision is what makes this paper so exciting. Let's get into it.

First Page: Tom: We're back with "Generating particle physics Lagrangians with transformers." Jane, we've talked about the results and the fixes. But the first page of this paper is where they lay out the big dream.

Jane: It's ambitious. They want to build a foundation model for theoretical physics. Not just to generate Lagrangians, but to identify, process, and manipulate equations. They see this as the first step toward a tool that could explore physics beyond the Standard Model.

Lu: And that's the part that gets me. The Standard Model is incomplete. We know there's dark matter, we know there are neutrino masses. But we don't know what the Lagrangian looks like. A model like this could help physicists explore the space of possible theories.

Meng: But it's a long way from generating Lagrangians to discovering new physics. The paper is clear that this is a proof of concept. The real value is in showing that transformers can handle this kind of symbolic, symmetry-rich mathematics.

Tom: And they make a great comparison. In natural language, words have context. In physics, fields have quantum numbers. The transformer's attention mechanism can capture those relationships.

Jane: Exactly. They draw a parallel between grammar and symmetry. A Lagrangian has to obey the rules of symmetry, just like a sentence has to obey the rules of grammar. The model learns those rules.

Lu: The embedding analysis is fascinating. They show that the model groups fields by their representations. Scalars cluster together, fermions cluster together. And they even found a "conjugation axis"—a consistent direction in the embedding space that represents the operation of taking a field to its antiparticle.

Meng: That’s the kind of interpretability we need. It’s not just a black box that gives the right answer. We can see that it’s learned the abstract structure of the physics.

Tom: So the first page sets the stage for a much bigger journey. It’s not just about this paper; it’s about the future of AI in theoretical physics.

Jane: And that future is what we’ll wrap up with in our final segment. Don't go anywhere.

Conclusion: Tom: And we're back for the final word on "Generating particle physics Lagrangians with transformers." Jane, it's been a wild ride. Let's bring it all together.

Jane: It really has. The paper shows that a transformer model can learn to write Lagrangians—the equations that describe particle physics—with over ninety percent accuracy. It’s not perfect, but it’s a massive step forward.

Lu: The key takeaway for me is the generalization. The model doesn't just memorize; it understands the symmetries. It can handle out-of-distribution scenarios, even if it struggles with the details. That’s the sign of a true foundation model in the making.

Meng: And from an engineering perspective, the failure modes are clear. Counting issues, rare term types, and number representation. These are all addressable with better architecture and data. The path forward is concrete.

Tom: And the authors are generous with their work. The model and datasets are public. Anyone can build on this. That’s how we get from a proof of concept to a tool that actually helps physicists discover new physics.

Jane: The ultimate goal is ambitious—to connect experimental data to theoretical models. Imagine feeding the model data from the Large Hadron Collider and having it suggest Lagrangians that explain the anomalies. That’s the dream.

Lu: And it's not a pipe dream. This paper shows the foundation is solid. The next steps are about scaling up and refining.

Tom: So we say goodbye to this paper, but not to the ideas. It's a stepping stone, and we can't wait to see what comes next.

Jane: Thanks for joining us, everyone. We'll be back soon with more exciting research from the arXiv. Until then, keep asking big questions.

Tom: And keep looking at the stars. See you next time!

Yong Sheng Koay, Rikard Enberg, Stefano Moretti, Eliel Camargo-Molina

Uppsala University · University of Southampton

cs.LG, cs.SC, hep-ph, hep-th

Submitted: 2025-01-16

Comments: 32 pages, 11 figues, 18 tables

Journal ref: SciPost Phys. 21, 024 (2026)

DOI: 10.21468/SciPostPhys.21.1.024

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 69/100

The gist: This paper presents a novel application of transformer models to generate particle physics Lagrangians from a given list of particle fields.

Key concepts

Lagrangian
In simple terms, a Lagrangian is the instruction manual for a physical system. It is an equation that tells physicists how particles move and interact within the universe.
Transformers
This refers to a type of AI architecture (like those powering ChatGPT) trained to process sequences. Here, it was used to automatically write complex equations (Lagrangians) instead of requiring human expertise.
Standard Model
The Standard Model is the current theory describing fundamental particles and forces. The paper tested the AI's ability to generate Lagrangians that adhere to the specific symmetries of this model.
Symmetries
These are mathematical rules about how particles transform. For a Lagrangian to be physically accurate, it must obey these rules, which the AI was shown it could learn.

Terminology

Summary

This paper presents a novel application of transformer models to generate particle physics Lagrangians from a given list of particle fields. The authors trained a Bidirectional and Auto-Regressive Transformer (BART) architecture with approximately 357 million parameters to predict Lagrangians respecting the Standard Model SU(3) × SU(2) × U(1) gauge symmetries. The work is framed as a step forward in the application of transformer models in symbolic scientific work, going beyond previous efforts that tackled less complex mathematical expressions such as arithmetic, linear algebra, or calculus.

The authors draw parallels between linguistic structure and Lagrangian structure: fields, their combinations (terms) and the invariance of the whole expression under symmetry resemble words, sentences and grammatical structures, respectively. They emphasize that a single gauged fermion field symbol ψL contains more information than a single variable x and a covariant derivative symbol Dµ contains more information than a single differential operator d/dx.

The stated long-term objective is to build a foundation model capable of identifying, processing, manipulating and exploring equations, indeed, a basis for theoretical physics. A medium-term goal is a transformer-powered system to explore theories Beyond the SM (BSM) of particle physics.

The authors built a pipeline combining AutoEFT and custom code to generate training data. AutoEFT finds allowed interaction terms, while custom code generates kinetic and mass terms. The dataset focuses on the SU(3) × SU(2) × U(1) gauge group, considering scalar fields (spin-0) and fermion fields (spin-1/2). Fields are characterized by their representations under each gauge group: for SU(3), fields can be in triplet (3), antitriplet (3̄), or singlet (1) representations; for SU(2), singlet (1), doublet (2), or triplet (3); for U(1), hypercharges are assigned as random fractions with numerators from −9 to 9 and denominators from 1 to 9, simplified to lowest terms.

Two datasets were created: a uniform dataset and a sampled dataset of approximately 280K Lagrangians. The sampled dataset used a probability distribution favoring fewer fields: 25% chance each for Lagrangians to have one, two, or three fields, an 11% chance for four fields, and a 7% chance each for five or six fields. The dataset was enriched so that approximately 50% of the Lagrangians with more than two fields contained trilinear interactions. The authors note this approach was inspired by findings that there are significant gains when optimizing the dataset for learning and that train set priming suggests that the opposite might be true regarding longer examples being harder to learn.

A custom tokenization scheme was developed with a full vocabulary including tokens for fields, spin, helicity, dagger, derivatives, sigma bars, symmetry groups, contractions, and identifiers. Fields are tokenized with the order: FIELD token, spin tokens, symmetry tokens, helicity tokens (if required), dagger token (if required), and ID token. Covariant derivatives implicitly encode gauge fields, representing a trade-off that reduces token count without sacrificing essential physics. Contraction information is explicitly provided after a CONTRACTION token, specifying how indices are contracted.

Both models achieved high accuracy on test datasets. On the merged test dataset, the sampled model achieved 92.0% perfect Lagrangian scores, while the uniform model achieved 93.2%. The authors report: Each model works relatively well on their own dataset while the uniform model excels in minimizing length penalization. The sampled model performed better on object and contraction scores (less than 5% errors), while the uniform model performed slightly better on length (less than 1% having extra terms).

The authors note: In almost every case (except for the 6-field scenario by the sampled model), close to 99% of the predicted Lagrangians have less than 20% errors. They also observed that "even on expected Lagrangians from the training dataset, the model generates Lagrangians with ID tokens that differ from the expected Lagrangians. This strongly suggests that the model has learned the concept of dummy indices."

The authors tested two OOD scenarios: higher numbers of input fields (up to 10, beyond the 6-field training limit) and unseen U(1) hypercharges. For higher field counts, more than 99% of the predicted Lagrangians are reasonable even in OOD scenarios, though there is a steady decrease in predicted Lagrangians that conserve U(1) symmetry as it goes towards more OOD scenarios. The model's performance decreases with more input fields, with the main failure mode being missing terms: the model is mainly missing terms rather than consistently predicting the wrong terms.

The authors attribute these failures to architectural limitations: Non-causal Transformers, like BART's bidirectional encoder, are known to struggle with contextual counting tasks. They note the model fails to handle more than 7-8 fields, struggling to count the 'FIELD' tokens in the input.

For OOD U(1) charges, the model's performance degrades with more OOD fields and input fields. The authors note: In OOD Lagrangians with trilinears, it is likely that the model memorized possible charge combinations that would lead to trilinear terms instead of learning how to sum up charges to conserve U(1).

The authors performed t-SNE visualization of field embeddings, finding distinct clusters that separates scalars and fermions and that singlet fields tend to occupy a central region within the t-SNE space, while gauged fields are distributed towards the sides. They also investigated whether the model learned conjugation as a consistent vector offset in embedding space, finding "the conjugation vectors have a preferred axis, as indicated by the peak near 1, whereas the randomly translated vectors are evenly distributed. This suggests that the model has learned the concept of conjugation as a specific transformation in the embedding space."

The authors evaluated the model on existing particle physics models. For simpler models like Scalar QED and the Lepton Sector, the model achieved perfect scores. Performance decreased with more complex models: the Standard Model achieved a best Lagrangian score of 0.77, with the main failure mode being Missing Yukawas. The authors note: "the set of fields in the Standard Model features the unique property of having six Yukawa terms. As a result, the model struggles to generate the one-generation Standard Model as it tends to miss four out of the six Yukawa terms."

The authors conclude: We have shown that, given a list of particle fields, transformer models are capable of writing the corresponding particle physics Lagrangian. They note that the concepts of symmetry groups and representations, conjugation operation and even the use of dummy indices, have all been learned by the model.

Future work will investigate adjusted transformer architectures or enhanced datasets to tackle these limitations, aiming to scale the approach toward building a broader 'foundation model' for particle physics. The authors envision to eventually input into the ML environment aimed at constructing theoretical Lagrangians also experimental information, in the form of real data showing deviations from the SM predictions.

The trained models are available at https://huggingface.co/JoseEliel/BART-Lagrangian, training datasets at https://huggingface.co/datasets/JoseEliel/lagrangian generation, and an interactive demonstration at https://huggingface.co/spaces/JoseEliel/generate-lagrangians.

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems and what the improved systems can do:

  • Improvement: Extend transformer architectures to handle symmetry-rich mathematical objects (fields with quantum numbers, covariant derivatives, gauge groups) rather than just simple variables and operators.

  • Capability: The improved system can generate physically valid Lagrangians for particle physics models, respecting SU(3)×SU(2)×U(1) gauge symmetries with over 90% accuracy.

  • Improvement: Implement a custom tokenization scheme that encodes quantum numbers (spin, color charge, weak isospin, hypercharge) as structured token sequences, including contraction information for group indices.

  • Capability: The system can parse and generate complex physics expressions while preserving all symmetry information needed for validation.

  • Improvement: Address the model's inability to count tokens in out-of-distribution scenarios (beyond 6 input fields) by incorporating counting-aware attention mechanisms or auxiliary counting objectives.

  • Capability: The improved system maintains >99% reasonable output generation even with 10 input fields, though accuracy decreases from 95% to 87% for U(1) conservation.

  • Improvement: Replace text-based number tokenization with continuous numerical encoding (as suggested by the paper's analysis of U(1) charge arithmetic failures) to improve arithmetic reasoning.

  • Capability: The system can correctly compute charge conservation in trilinear interactions with OOD fractional charges, where current text-based encoding drops accuracy to 20%.

  • Improvement: Leverage the discovered embedding structure (symmetry clusters, conjugation axes) to build interpretable representations that can be probed for physical properties.

  • Capability: The system can identify conjugation relationships between fields via consistent vector offsets in embedding space, enabling automated symmetry analysis.

  • Improvement: Implement the paper's finding that oversampling shorter sequences (25% each for 1-3 fields) improves performance on longer sequences, using a curriculum that starts with simple cases.

  • Capability: The system achieves comparable or better performance on 6-field Lagrangians while training on only 7% such examples, versus 16% in uniform datasets.

  • Improvement: Add targeted enrichment for trilinear interactions (scalar trilinears and Yukawa terms) during dataset generation, as these are rare in uniform sampling but crucial for physics.

  • Capability: The system correctly generates trilinear interactions with >97% accuracy on in-distribution data, versus 90% without enrichment.

  • Improvement: Integrate physics validation (mass dimension checks, U(1) conservation, contraction consistency) into the generation loop as a post-processing filter.

  • Capability: The system can self-correct outputs, ensuring >99% of generated Lagrangians are reasonable even in OOD scenarios.

  1. Automated Theory Building: Given a list of particles with quantum numbers, generate complete, physically valid Lagrangians including kinetic, mass, and interaction terms.

  2. BSM Exploration: Systematically explore beyond-Standard-Model theories by generating Lagrangians for novel field content, enabling rapid hypothesis generation.

  3. Symmetry Verification: Automatically verify that generated expressions respect all imposed gauge symmetries, with explicit contraction information.

  4. Interpretable Physics Representations: Provide embeddings that encode physical concepts (group representations, conjugation) in a structured, analyzable way.

  5. OOD Robustness: Handle unseen field counts and unusual hypercharges with graceful degradation, maintaining structural validity even when exact accuracy decreases.

  6. Foundation Model for Theoretical Physics: Serve as a basis for future systems that can process, manipulate, and explore equations across particle physics, potentially integrating experimental data for data-driven discovery.

Sources

Related papers