Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics

arXiv:2606.26228 · physics.data-an, astro-ph.GA, cs.LG, hep-ph · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: Next we'll be talking about the paper "Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics".

Jocelyn: The paper was written by Rikab Gambhir, Luisa Lucie-Smith and Jesse Thaler from University of Cincinnati and Universität Hamburg and Massachusetts Institute of Technology and Institut des Hautes Études Scientifiques and CEA Paris-Saclay and The NSF Institute for Artificial Intelligence and Fundamental Interactions.

Vera: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 1: Vera: We are moving from our look at specific cosmic observations into a more foundational discussion about how we actually process information in modern physics. We are looking at a very important new piece titled 'Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics' by Rikab Gambhir, Luisa Lucie-Smith, and Jesse Thaler. It is part of this larger VERaiPHY initiative which is basically trying to set some ground rules for how we use AI in things like high energy physics and cosmology.

Jocelyn: It sounds like they are finally addressing the elephant in the room, doesn't it? For a long time, people have just been throwing these massive neural networks at data and hoping for the best.

Vera: Exactly, and these authors are coming from some heavy-hitting institutions like MIT and the University of Cincinnati to say we can't just treat these models as magic boxes anymore. They are arguing that if we want to do real science, we have to stop using these terms like "transparency" or "trustworthiness" interchangeably because they mean very different things when you are trying to understand a galaxy or a particle collision.

Subrahmanyan: I think it is crucial that they are bringing this up now, especially as these models get larger and more complex. If we cannot distinguish between the structure of the math and the physical meaning of the result, we are essentially doing statistics without any connection to reality. It is one thing to have a model that predicts a number correctly, but it is quite another to know if that number actually tells us something about the universe.

Jocelyn: That distinction between the math and the physics seems to be their main starting point. They aren't just talking about computer science; they are talking about the philosophy of how a physicist works.

Subrahmanyan: Precisely, Jocelyn. They are suggesting that machine learning models should be held to the same rigorous standards as our classical models, like General Relativity or the Standard Model. The only real difference is the scale and complexity, not the fundamental requirement for understanding.

Vera: It is a bold claim to say that a neural network is just a "classical" model in a different skin, but it sets the stage for how they define their terms. Let's look at how they actually draw those lines between interpretability and explainability.

Paper discussion segment 2: Jocelyn: Now that we have the context of who is writing this, let's get into what 'Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics' actually proposes as its core framework. They make this very sharp distinction that I think will change how researchers approach their model design.

Vera: Right, they define interpretability as being about the structural transparency of the model itself. It is basically asking if we can look at the inner workings, the actual computation or the "mechanics," and understand how an input turns into an output. It is a question of form and structure rather than what that structure actually represents in our physical world.

Subrahmanyan: And then there is explainability, which they treat as a completely different axis. Explainability isn't about the math; it's about whether we can map those mathematical results onto our existing domain knowledge, like laws of gravity or particle interactions. You could have a model that is perfectly interpretable because you can see every single multiplication happening, but if those multiplications don't correspond to a physical concept, then it isn't explainable at all.

Jocelyn: That is such an important point because it explains why we often feel so lost with deep learning. We might be able to see the code and the weights, which makes it somewhat interpretable in a technical sense, but we have no idea how those weights relate to, say, the density of dark matter.

Vera: They even use a quadrant system to show that you can have one without the other. You could have something that is explainable but not interpretable—like a complex mathematical function where we know the physical result but the actual numerical computation is a total black box.

Subrahmanyan: Or you can have the most dangerous case: something that is neither. If a model is both opaque in its structure and disconnected from physical principles, it provides no scientific value; it's just a tool for prediction without any capacity for insight.

Jocelyn: It really forces you to realize that "understanding" isn't just one single thing. Let's talk about the trade-offs they mention, because you can't just have everything at once.

Paper discussion segment 3: Vera: We are digging into the practical reality of these definitions now, specifically looking at the costs associated with them in 'Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics'. The authors point out that there is no free lunch here; every choice you make has a direct consequence for your science.

Jocelyn: They frame it as a tension between interpretability and expressivity, and then between explainability and adaptability. If you want a model that is incredibly easy to interpret, like a simple linear regression or a basic parametric fit, you are going to lose the ability to capture the incredibly complex, messy patterns found in real cosmic data.

Subrahmanyan: That is the fundamental struggle of modeling. If you constrain your model too much so that it's easy for a human to read, you might miss the very phenomena you are looking for because your model was too simple to "express" them. On the other hand, if you let a neural network be as expressive as it wants, it will fit the data perfectly, but it will be an uninterpretable mess that tells you nothing about the underlying physics.

Vera: And then they bring up adaptability, which is tied to explainability. If you force a model to adhere strictly to known physical laws—like making sure a model respects Lorentz invariance—you are increasing its explainability, but you might be limiting its ability to adapt to new data that suggests our current laws are incomplete.

Jocelyn: It's like the difference between a rigid theory and an empirical fit. One is deeply connected to what we know, but it might be too stiff; the other is flexible, but it might just be memorizing noise.

Subrahmanyan: They also suggest that we need to be much more intentional about our "intervention plans." It isn't enough to just run a post-hoc test like a Shapley value and say, "look, this feature is important." You have to decide beforehand what you would actually do with that information if it turned out to be unexpected.

Vera: That's a great point about the methodology. They suggest moving toward "intrinsic" methods where we build the physics directly into the architecture, rather than just trying to explain it after the fact.

Conclusion: Vera: We have covered a lot of ground today, from the basic definitions to the difficult trade-offs that every physicist using AI has to face. This paper really pushes us to be more disciplined in how we design our experiments and our models.

Jocelyn: It really does. It's not just about getting a better accuracy score on a leaderboard; it's about ensuring that the work we do actually contributes to our collective understanding of the cosmos.

Subrahmanyan: I think the most vital takeaway is that machine learning isn't a separate branch of science that operates by different rules. It is simply a new way of building models, and it must be integrated into our existing scientific framework with all the same skepticism and rigor we apply to any other theory.

Vera: Well said. We've spent our time today looking at 'Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics', and it has certainly given us a new lens through which to view the data coming in from our telescopes and colliders.

Jocelyn: It's a heavy topic, but a necessary one if we want to turn these black boxes into true instruments of discovery.

Subrahmanyan: I agree; the era of treating AI as a magic oracle must end if we are to enter the era of machine-driven discovery.

Vera: Thank you all for listening. We'll be back next time with another look at what's happening on the frontiers of research. Goodbye for now.

Jocelyn: Goodbye!

Subrahmanyan: Goodbye everyone.

Rikab Gambhir, Luisa Lucie-Smith, Jesse Thaler

University of Cincinnati · Universität Hamburg · Massachusetts Institute of Technology · Institut des Hautes Études Scientifiques · CEA Paris-Saclay · The NSF Institute for Artificial Intelligence and Fundamental Interactions

physics.data-an, astro-ph.GA, cs.LG, hep-ph

Submitted: 2026-08-16

Updated: 2026-08-18

Comments: 31 pages, 3 figures, Part of the VERaiPHY Initiative

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: This paper reviews the concepts of interpretability and explainability as they apply to machine learning in physics.

Key concepts

Interpretability
This refers to the structural transparency of a model, asking if one can examine the inner workings or computation to understand how an input leads to an output. It is about form and structure rather than what that structure physically represents.
Explainability
This is about mapping mathematical results onto existing domain knowledge, such as laws of gravity or particle interactions. A model can be interpretable structurally but not explainable if its computations do not correspond to a physical concept.
Interpretability vs. Explainability Trade-off
There is a tension between these concepts and other factors like expressivity and adaptability. Constraining a model for easy interpretation might limit its ability to capture complex data, while high expressivity can lead to uninterpretable results.

Terminology

Summary

This paper reviews the concepts of interpretability and explainability as they apply to machine learning in physics. The authors define interpretability as concerning the structural transparency of a model—the ability to understand, or approximate, the inner workings of a model and how it reaches its output—and explainability as concerning the scientific content of a model—the ability to map the model onto existing knowledge in the relevant scientific domain. They emphasize that ML models are no different than classical models, except in size, and that interpretability and explainability are best understood as deliberate modeling choices rather than inherent properties.

The paper discusses the trade-offs each entails: interpretability vs. expressivity and explainability vs. adaptability. It explores the contexts in which each is needed, and the intrinsic and post-hoc tools available for achieving them. The authors state that interpretability is primarily a feature of the form of the model f and its dependence on θ and x, not of the specific values of θ or fθ(x), while explainability does not exist without domain knowledge to compare to.

The paper emphasizes the importance of task specification and intervention plans as a core aspect of model design. It argues that interpretability and explainability are not always necessary, and this determination is highly problem dependent and subjective. The authors also stress that one can design an ML model, loss function, or training technique in a way that aligns with their desired scientific goals, which may or may not include interpretability and explainability.

The paper is organized as follows: Section 2 provides foundations and core definitions; Section 3 focuses on interpretability; Section 4 focuses on explainability; Section 5 discusses intrinsic and post-hoc methods; Section 6 discusses task specification and intervention planning; and Section 7 concludes. The paper contributes to the VERaiPHY initiative, a PHYSTAT review series establishing verification and validation standards for ML across particle physics, astrophysics, and cosmology.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems, along with what the improved system can do:


Improvement: Add two distinct, measurable evaluation axes to any AI system:

  • Interpretability Score (structural transparency): How well can the system’s internal computations be approximated or read off (e.g., via decision paths, linear coefficients, or distillation error)?

  • Explainability Score (domain-knowledge alignment): How well do the system’s learned parameters or features map onto established physical/scientific concepts (e.g., mass, force, symmetry)?

What the improved system can do:

  • Automatically report both scores after training, allowing users to see if a model is “interpretable but not explainable” (e.g., a neural net with clear activations but no physical meaning) or vice versa.

  • Flag models that are neither, preventing deployment of scientifically useless black boxes.

Improvement: Before training, require the user to specify:

  • The downstream scientific question (not just the loss function).

  • The desired level of interpretability/explainability (low, medium, high).

  • A pre-defined intervention plan: What specific SHAP value, saliency map, or latent variable behavior would cause the user to reject, modify, or trust the model?

Improvement: For physics applications, automatically inject known symmetries (e.g., Lorentz invariance, rotational equivariance) or conservation laws into the architecture or loss function, rather than relying on the network to learn them.

Improvement: Add a built-in post-hoc distillation step that approximates any trained black-box model with a simpler, interpretable surrogate (e.g., a symbolic expression via PySR or a shallow decision tree), and report the approximation error.

Improvement: For deep networks, add an optional sparse autoencoder layer that decomposes the model’s internal representations into a small set of interpretable “atoms” or concepts.

Improvement: Automatically detect if the model is overfitting by memorizing the training data (e.g., behaving like a look-up table) rather than learning generalizable patterns.

Improvement: Provide tools to map learned latent variables or feature importances onto known physical quantities, and flag any novel, unexplained correlations.

  • Be transparent: Users can always see how the model works (interpretability) and what it means (explainability).

  • Be trustworthy: Pre-defined intervention plans prevent confirmation bias and ensure that any interpretation leads to a concrete action.

  • Be scientifically useful: Models are constrained by physics principles, so they generalize better and can be used for discovery, not just prediction.

  • Be adaptable: Users can choose the trade-off between interpretability and expressivity based on their specific task, rather than being forced into one extreme.

  • Be safe: The system will refuse to deploy models that are neither interpretable nor explainable, or that behave like look-up tables, unless explicitly overridden.

Abstract

We review the concepts of interpretability and explainability as they apply to machine learning in physics. We define interpretability as concerning the structural transparency of a model (the ability to understand or approximate its inner workings) and explainability as concerning the scientific content of a model (the ability to map it onto domain knowledge). We discuss the trade-offs each entails (interpretability vs. expressivity; explainability vs. adaptability), the contexts in which each is needed, and the intrinsic and post-hoc tools available for achieving them. Throughout, we emphasize that machine-learned models are subject to the same scientific questions as classical models, differing only in scale, and that interpretability and explainability are best understood as deliberate modeling choices rather than inherent properties. We also emphasize the importance of task specification and intervention plans as a core aspect of model design.

Sources

Related papers