Beyond L 2: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

arXiv:2608.16773 · cs.LG · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Beyond L 2".

Tom: Prototype-based neural networks are hailed as interpretable-by-design architectures, and Abductive Latent Explanations (ALE) were introduced to provide formal,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, the paper explains that prototype-based neural networks are considered interpretable-by-design architectures, but the old ALE methods were too narrow for these newer designs; they introduce this generalized framework to provide formal, mathematically guaranteed explanations by leveraging the network's intrinsic structure.

Jane: In simple terms, they’re taking a powerful explanation technique and making it adaptable so it can work even when the latent space isn't a simple flat Euclidean space. It’s about bridging that gap between theory and modern practice.

Lu: Specifically, they look at how prototypes are activated differently in various architectures; for example, ProtoPNet starts by computing the L2-distances d l,j = z l - p j squared between each patch z l and each prototype p j, and then defines the similarity sim (z l, p j) as d l,j squared + epsilon l,j where epsilon is a small constant.

Meng: The paper details how they handle these different activation measures by deriving specific bounds for each case—whether it’s bounding the log distance squared or dealing with angular distances in spherical spaces. I need to see exactly how they manage those different math types.

Lalam: It seems like the central mechanism involves computing tight bounds on latent space distances to produce those formal explanations, but now those bounds are being extended beyond just Euclidean geometry into more complex realms.

Tom: And they validate this whole approach by computing subset-minimal formal explanations on fully trained image classifiers, which is a big deal for proving the method actually works in practice and isn't just theoretical fluff.

Jane: The key finding here is that by unifying these diverse models under a single formal framework, they enable the first rigorous, cross-architecture comparison of their interpretability methods across different network designs.

Lu: This means we are moving from just observing one network type to actually understanding the fundamental limits of explainability across all prototype-based designs they tested.

Meng: That’s a significant methodological shift because it validates that formal guarantees can exist even when the underlying mathematical structure of the AI is highly varied and complex.

Lalam: For me, this advance suggests that we are finally moving toward a standardized way to audit AI decisions, which is huge for building trust in critical applications where reliability matters most.

The paper's summary: Tom: What really stands out in this paper is how they tackle the different architectural variants head-on; they don't just assume everything fits one mold; they systematically address top-k explanations, cosine similarity models, and dimensional projection methods separately.

Lu: For example, when dealing with Cosine Similarity models like TesNet, which operate on spherical latent spaces where vectors are normalized, they derive bounds using angular distance and a technique called Cosine Similarity Bounds via Monotone Inversion.

Meng: That sounds mathematically intensive; it suggests they’ve developed specific geometric reasoning techniques to translate those non-Euclidean similarities back into something useful for bounding the activation values.

Lalam: It’s like they built a translator that converts the complex geometry of these new architectures into the language of formal distance constraints we need for verification.

Jane: They also handle architectures using softmax functions over prototypes, like PIP-Net, by focusing on the conservation of probability mass to define similarity bounds within a zero one range.

Tom: That’s a really smart move because it shows they aren't just sticking to one style; they are creating architecture-specific algorithms for every geometric variant we see in the field.

Lu: The implication is that we can now systematically derive how to bound activations for any prototype-based network, which is a huge leap in the field of formal verification.

Meng: Practically speaking, this means developers can start choosing their network structure based on how well it supports formal verification from the get-go, which is a huge shift in model design philosophy.

Lalam: For me, this advance means we can build a culture where rigorous explanation isn't an afterthought but something inherently built into the architecture itself for every AI system we deploy.

The paper's improvements: Tom: So, to wrap things up, the main point of "Beyond L2: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures" is that they’ve successfully generalized ALE to handle non-Euclidean prototype architectures by constructing novel bounding algorithms for spherical metrics and dimensional projections.

Jane: Essentially, they took the concept of formal explanation and made it flexible enough for modern, diverse AI models, moving us away from the old Euclidean constraint and into a much more practical realm for real deployment.

Lu: This opens up a massive avenue where we can rigorously compare the interpretability of completely different classes of prototype networks using these new generalized bounds that were previously out of reach.

Meng: Practically speaking, this means developers can start choosing their network structure based on how well it supports formal verification from the get-go, which is a huge shift in model design philosophy for building AI.

Lalam: For me, this advance is huge because it means we can build a culture where rigorous explanation isn't an afterthought but something inherently built into the architecture itself for every AI system we deploy.

Tom: It’s a fantastic paper that gives us concrete tools to move from abstract interpretability promises to actual, verifiable safety guarantees for complex models.

Jane: Absolutely, and I think this paper will fundamentally change how we approach XAI by making formal verification accessible across more model types.

Lu: That’s the big picture here; we can start comparing different AI designs with mathematical rigor that was previously out of reach for us to achieve.

Meng: I just hope these tools translate into real-world deployment speed improvements, because theoretical guarantees are great, but they still have to run efficiently on hardware for this to be truly useful.

Lalam: That’s the vision we need; a culture where rigorous explanation isn't an afterthought but something inherently built into the architecture itself for every AI system we deploy.

Conclusion: ---: Conclusion ---

Tom: So, to wrap up our discussion on "Beyond L2: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures," we’ve seen how this research successfully generalized ALE to handle a much wider variety of prototype architectures than before.

Jane: Exactly, Tom; they took that powerful explanation technique and made it flexible enough for modern AI systems that aren't just using simple flat Euclidean spaces anymore.

Lu: This opens up a massive avenue where we can rigorously compare the interpretability of completely different classes of prototype networks using these new generalized bounds that were previously out of reach.

Meng: Practically speaking, this means developers can start choosing their network structure based on how well it supports formal verification from the get-go, which is a huge shift in model design philosophy.

Lalam: For me, this advance is huge because it means we can build a culture where rigorous explanation isn't an afterthought but something inherently built into the architecture itself for every AI system we deploy.

Tom: It’s truly a fantastic paper that gives us concrete tools to move from abstract interpretability promises to actual, verifiable safety guarantees for complex models.

Jane: Absolutely, and I think this paper will fundamentally change how we approach XAI by making formal verification accessible across more model types.

Lu: That’s the big picture here; we can start comparing different AI designs with mathematical rigor that was previously out of reach.

Meng: I just hope these tools translate into real-world deployment speed improvements, because theoretical guarantees are great, but they still have to run efficiently on hardware.

Jules Soria, Alban Grastien, Romain Xu-Darme, Julien Girard-Satabin, Zakaria Chihani, Daniela Cancila

Université Paris-Saclay

cs.LG

Submitted: 2026-08-17

Updated: 2026-08-17

Code: https://github.com/julsoria/beyond_l2

Importance score: 83/100

The gist: Prototype-based neural networks are hailed as interpretable-by-design architectures, and Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations

Key concepts

Abductive Latent Explanations (ALE)
A technique introduced to provide formal explanations for AI models. The paper generalizes ALE to handle diverse prototype-based architectures, making the explanation framework adaptable even when the latent space is not a simple flat Euclidean space.
Prototype-based neural networks
Architectures considered interpretable-by-design because they use prototypes. The research focuses on how these prototypes are activated differently across various network designs, such as ProtoPNet.
Generalized Bounds
Novel algorithms derived to compute tight bounds on latent space distances for non-Euclidean metrics, including angular distances and dimensional projections. These bounds allow the method to work across different mathematical structures.
Formal Verification
The process of using mathematical rigor to prove that AI explanations are accurate and reliable. The paper provides concrete tools that move interpretability from abstract promises to verifiable safety guarantees for complex models.

Terminology

Summary

Prototype-based neural networks are hailed as interpretable-by-design architectures, and Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic structure of these networks to ensure both predictive safety and human readability. ALEs rely on computing tight bounds on latent space distances to produce formal explanations. However, existing ALE formulations are rigidly confined to Euclidean latent spaces. This leaves a critical gap: modern state-of-the-art architectures increasingly rely on non-Euclidean representations—spherical metrics, Gaussian densities, and dimensional projections—rendering current formal explanation methods incompatible. In this work, the authors generalize the ALE framework to support non-Euclidean prototype architectures. For each geometric variant, they systematically derive how to either map the architecture to existing bounds or construct novel, architecture-specific bounding algorithms. They validate their theoretical constructions by computing subset-minimal formal explanations on fully trained image classifiers. By unifying these diverse models under a single formal framework, they enable the first rigorous, cross-architecture comparison of their interpretability.

The work is performed in the context of prototypical part networks that classify images into one of C classes. The input image X is passed through a neural network backbone f that produces a latent representation z = f (X) ∈ Z where Z = RH1 ×W1 ×D. This latent representation z is interpreted as a collection of (H1 ×W1) vectors zl ∈ RD (or patches) where l ∈ L = L x L. In the seminal work, ProtoPNet [9] starts by computing the L2-distances dl,j = ∥zl − pj ∥2 between each patch zl and each prototype pj, and then defines the similarity sim(zl, pj) as log dl,j squared +ϵ l,j. Finally, the activation aj of prototype pj is the maximum similarity over all patches of the image: aj = maxl∈L sim(zl, pj). Regardless of how the activation vector a is computed, the last step feeds the activation vector to a linear layer W of dimension P × C that attributes each prototype a contributing weight for each class. The classification is made by selecting the class that has the highest score: argmax W T a.

An ALE [40] is defined as a collection E of statements about the latent representation of the input image, together with an implicit interpretation of this collection that allows one to infer some information about the classification of the image. Such statements are made in relation with existing prototypes, e.g., the degree of similarity between a patch and a prototype. ALEs allow one to bound similarity or activation values: sim(zl, pj) ∈ [simE (zl, pj), simE (zl, pj)] and aj ∈ [aE,j, aE,j]. When the exact value for a similarity or an activation is reported in the explanation, this value explicitly trivializes the interval: (zl, pj) ∈ E ⇒ simE (zl, pj) = sim(zl, pj). Given these intervals, it is then possible to give bounds on the difference between the scores of each class. Generating an ALE is an iterative process where adding a prototype—or adding a ⟨patch, prototype⟩ pair—to the explanation E tightens the bounds on the activation of other prototypes, and consequently, on the output logits of the model. If, in this way, a single class is proved to have a greater value than all other classes, the decision of the network is guaranteed (explained) and we say that the explanation E is formal.

The work addresses different geometric variants:

  1. Top-k Explanations: These provide k most activated prototypes, guaranteeing that every prototype not included in the ALE has an activation lower than those included. The algorithm incrementally adds the prototype whose activation is highest amongst those not yet included until the model’s prediction is guaranteed by the explanation.

  2. Cosine Similarity: Models like TesNet rely on dot-product similarity sim(zl, pj) = zl · pj, implying a spherical latent space where vectors are normalized (∥zl∥ = 1). Since cosine similarity is not a distance metric and violates the triangle inequality, the authors derive bounds using angular distance d∠ (x, y) and Cosine Similarity Bounds via Monotone Inversion: simE (zl, pk) = cos min π, d∠ (pk, pint) + rint, where pint and rint are derived from a Spherical Cap Intersection Approximation.

  3. Dimensional Projection: For PIP-Net employing a softmax function over prototypes, the similarity sim(zl, pj) is defined as ẑ l,j = softmax(zl), meaning it is a conservation of probability mass with sim(zl, pj) ∈ [0, 1], and j∈P sim(zl, pj) = 1.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to AI systems by implementing the generalized Abductive Latent Explanation (ALE) framework:


The core improvement is shifting from rigid Euclidean distance constraints in explainability methods to a flexible, geometry-aware framework capable of handling diverse prototype-based architectures.

Here are the specific improvements and what they enable:

  1. ​​Improve Interpretability Across Diverse Architectures (Generalization):

  2. ​​Support Non-Euclidean Latent Spaces: The system can now provide formal, mathematically guaranteed explanations for a wide range of modern prototype networks, including those using spherical metrics (cosine similarity), Gaussian densities, and dimensional projections, which were previously incompatible with standard ALE formulations.

  3. ​​Rigorous Formal Verification of Predictions: By computing subset-minimal formal explanations (ALE), the system can provide a mathematical guarantee that the model's prediction is correct based on the evidence provided by prototypes. This moves XAI beyond post-hoc analysis to a level of verification, crucial for safety-critical systems.

  4. ​​Quantify Architectural Inductive Biases: The framework allows researchers to rigorously dissect which architectural components (e.g., sparse linear heads in PIP-Net vs. dense heads in ProtoPNet) lead to better formal interpretability, guiding the design of inherently more transparent models for specific tasks.

  5. ​​Enable Novel Explanation Paradigms: The paper introduces and validates several new geometric reasoning techniques that can be integrated into the ALE process, such as:

  6. ​​ Spherical Cap Intersection Approximation (HIA) for Cosine Similarity: This allows the system to generate tight, closed-form similarity bounds even when prototypes are defined on a unit hypersphere, providing precise bounds on activations outside the explanation.

  7. ​​ Handle Probabilistic and Focal Pooling Architectures: The system can bound activation scores in architectures using softmax/simplex explanations (like PIP-Net) by tracking the conservation of probability mass, allowing for formal verification even when prototype activations are defined as relative similarities rather than absolute distances.

  8. ​​ Decouple Geometry from Model Forward Pass: The proposed pipeline maps similarity bounds into a universal Euclidean space, applies geometric intersection logic (HIA), and remaps the results back to the model's specific metric space. This allows for rapid computation of explanations without requiring external solvers or costly iterations during inference.

The improved AI system can achieve the following specific capabilities:

  1. ​​Provide Why Explanations with Mathematical Certainty: Instead of just showing which prototypes were most similar, the system can state, The prediction is class C because prototypes P1 and P5 have activations within range [X, Y], and the geometric constraints guarantee that no other prototype could possibly yield a higher score.

  2. ​​Ensure Safety in Critical Domains: For applications like autonomous driving or medical diagnosis (as per the paper's motivation), this system provides an audit trail where the explanation itself is formally sound, ensuring that decisions are based on verifiable evidence rather than opaque correlations.

  3. ​​Benchmark Model Design: The framework enables developers to choose model architectures based on their inherent interpretability potential; for example, favoring models with sparse linear heads if formal verification is a primary requirement.

  4. ​​Efficient Explanation Generation: By using closed-form geometric bounds (like Corollary 2), the system can generate subset-minimal explanations in time comparable to or faster than standard methods, making complex models more practically transparent.

Sources

Related papers