Similarity-Distance-Magnitude Activations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Similarity-Distance-Magnitude Activations".
Jane: We introduce a new activation function and estimator designed to decompose epistemic uncertainty in language models into interpretable signals: SIMILARITY, DISTANCE, and MAGNITUDE.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, shifting gears a bit, let’s look at who wrote this paper and what the title actually communicates about the work itself. The authors are working on making uncertainty estimates in language models more interpretable by introducing these three specific components: Similarity, Distance, and Magnitude.
Jane: Right; the authors are trying to give us a structured way to decompose predictive uncertainty into these three distinct parts. It’s like they’re breaking down a complex feeling of "how sure am I?" into concrete measurements we can analyze.
Lu: The core concept is viewing the network as a metric learner, which means we can measure its relationship to the training data set directly. This approach lets them derive signals about how well an input aligns with known examples during classification.
Meng: I’m interested in the mechanism they use to extract these signals; they’re using this exemplar adaptor to distill representations conditional on predictions. I wonder if that distillation process is computationally intensive enough for a high-throughput environment like ours.
Lalam: If we can distill the representation space based on what the model predicts, that could lead to much more nuanced safety guards for our AI systems because we’d know exactly where its confidence is coming from.
The paper's summary: Tom: So, to summarize "Similarity-Distance-Magnitude Activations," they introduce this SDM activation function, which modifies the standard softmax to be more robust and interpretable. It explicitly adds awareness about how similar an input is to training data and how far it is from the known distribution.
Jane: That’s the central idea, Tom; they define SIMILARITY as correctly predicted depth-matches into the training set, which tells us how well a new instance lines up with past successful examples. Then they layer on DISTANCE, which measures how far an instance is from its training distribution using class-wise empirical cumulative distribution functions derived from a labeled calibration set.
Lu: And then there’s MAGNITUDE, which captures the decision-boundary awareness by taking a value from the final linear layer of their exemplar adaptor. When you combine these three signals—Similarity, Distance, and Magnitude—they create a richer output distribution that sharpens where it should be sharp and reflects high uncertainty when things are far away.
Meng: So, if I get this right, the paper claims that combining these three signals makes the distribution sharper when q and d values are larger than in standard softmax. That suggests we get a better signal for selective classification under certain conditions.
Lalam: It sounds like they are providing a structured way to look at model confidence, moving beyond just saying "it's eighty percent sure" to telling us *why* it thinks so and how far off that might be in terms of its knowledge base.
The paper's improvements: Tom: The authors highlight the SDM estimator as a key improvement, which uses a data-driven partitioning of class-wise empirical CDFs over the SDM activation output to control accuracy among selective classifications. This is crucial for setting reliable thresholds.
Jane: They also show that this method is empirically more robust against covariate shifts and out-of-distribution inputs when compared to existing post-hoc calibration methods using standard softmax activations, while still keeping the results useful for in-distribution data.
Lu: They suggest defining a "HIGH-RELIABILITY (SDMHR) region" based on criteria like q' being greater than q' and the SDM output exceeding a certain threshold psi, which gives us a principled way to select the most trustworthy predictions.
Meng: That concept of the SDMHR region is what catches my eye for deployment; it lets us triage inputs—if it’s in that region, we treat it as reliable; if not, we flag it for human review or reject it outright. I really need to know how reliably q' functions when we deploy this in a live system.
Lalam: This triage capability is a massive step forward because instead of treating all AI outputs equally, this framework lets us automatically classify decisions into high-reliability, low-reliability, and rejected categories based on structural evidence rather than just relying on a simple probability score.
Conclusion: Tom: So to wrap up the discussion on "Similarity-Distance-Magnitude Activations," the main finding is that this framework provides a more structured way to quantify epistemic uncertainty by breaking it down into similarity, distance, and magnitude components. It offers a data-driven partitioning method for controlling how we manage accuracy among different selective classifications.
Jane: And they demonstrated that this approach handles shifts in the input distribution and out-of-distribution data more robustly than older calibration techniques while still being effective on the training data itself. Essentially, they've given us tools to check our AI outputs against the structure of what it was trained on.
Lu: The implication here is that we can build systems where we don’t just accept a prediction but actually understand the geometric relationship between an input and our knowledge manifold, which is a significant step for building more sophisticated AI architectures.
Meng: For practical application, this means we can create automated filtering pipelines that reliably reject inputs that are structurally too far from what the model has seen, which is vital for keeping systems stable and safe in operational environments.
Lalam: I think the most significant impact here is on building more trustworthy AI by providing concrete evidence of when an output is reliable enough to act upon, moving us toward a system that understands its own limitations.
Tom: That’s a fantastic summary of what we've discussed about the "Similarity-Distance-Magnitude Activations" paper. We’ve seen how this framework moves beyond simple probability scores into something much richer and more actionable for understanding model behavior.
Allen Schmaltz
cs.LG, cs.CL
Submitted: 2025-09-16
Updated: 2026-09-30
Comments: Published in Findings of the Association for Computational Linguistics: ACL 2026. 22 pages, 8 tables, 2 algorithms. (v6 adds Appendix A.10 and Alg. 2.) arXiv admin note: substantial text overlap with arXiv:2502.20167
Journal ref: Findings of the Association for Computational Linguistics: ACL 2026, pages 22037-22057
DOI: 10.18653/v1/2026.findings-acl.1109
Code: https://github.com/ReexpressAI/sdm_
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 79/100
The gist: We introduce a new activation function and estimator designed to decompose epistemic uncertainty in language models into interpretable signals: SIMILARITY, DISTANCE, and MAGNITUDE.
Key concepts
- SIMILARITY
- This signal measures how well a new input matches known examples from the training set. It is calculated by counting consecutive nearest matches that are correctly predicted and align with the true label of the test instance, indicating strong alignment with learned knowledge.
- DISTANCE
- Distance quantifies how far an input is from its training distribution. It is normalized using class-wise empirical Cumulative Distribution Functions (CDFs) derived from labeled data. This provides a principled measure of distributional novelty, showing how conservative the distance is relative to other instances in the same class.
- MAGNITUDE
- Magnitude represents the proximity of a prediction to its decision boundary. It is derived from the final linear layer of an exemplar adaptor. A larger magnitude suggests that the model's output is close to crossing a class threshold, reflecting high confidence or strong separation between classes.
Terminology
Summary
We introduce a new activation function and estimator designed to decompose epistemic uncertainty in language models into interpretable signals: SIMILARITY, DISTANCE, and MAGNITUDE. This method offers a robust alternative to standard softmax activations for selective classification, providing principled ways to control class- and prediction-conditional accuracy while enhancing robustness against covariate shifts and out-of-distribution inputs.
The gist
We introduce the SIMILARITY-DISTANCEMAGNITUDE (SDM) activation function, which is a more robust and interpretable formulation of the standard softmax activation function, adding SIMILARITY (i.e., correctly predicted depth-matches into training) awareness and DISTANCE-to-training-distribution awareness to the existing output MAGNITUDE (i.e., decision-boundary) awareness, and enabling interpretability-by-exemplar via dense matching.
How it works
The SDM activation function is defined as:
SDM(z′)i = (2 + q)d ·z′i for 1 ≤ i ≤ C (Equation 6). This function combines the three core signals derived from an exemplar adaptor: SIMILARITY, DISTANCE, and MAGNITUDE. The output distribution becomes sharper with larger values of q and d, as well as with larger relative values of z′yˆ,
similar to the standard softmax. When distance exceeds the largest observed distance in labeled data, d=0, resulting in a uniform distribution that reflects maximally high uncertainty.
The SIMILARITY (q) is defined as the count of consecutive nearest matches in Dtr that are correctly predicted and match yˆ of the test instance
(Equation 4). The DISTANCE (d) is normalized by defining it in terms of the class-wise empirical CDFs of dnearest over Dca, as the most conservative quantile relative to the distance to the nearest matches observed in the labeled, held-out set
(Equation 5). The MAGNITUDE is taken as z′yˆ,
derived from a final linear layer of an exemplar adaptor.
SDM Activation: Loss and Training
A corresponding negative log likelihood loss is introduced (Equation 7) that accounts for the change of base, which is necessary due to the modified activation function. This loss is used to train the weights of the exemplar adaptor, including its CNN and final linear layer parameters (G). The training process involves initializing with a standard softmax and then re-calculating q and d for each x ∈ Dtr after each epoch.
Evaluating Selective Classification
The paper introduces two key quantities for evaluating selective classifiers over a test set, Dte: (Quantity I) prediction-conditional accuracy at or above a given threshold, α, stratified by predicted labels (yˆ), and (Quantity II) class-conditional accuracy at or above the same threshold, α, stratified by true labels (y).
The core estimator is the SDM estimator based on a data-driven partitioning of the class-wise empirical cumulative distribution functions (eCDFs) over the SDM activation output.
This leads to defining a HIGH-RELIABILITY (SDMHR) region
using criteria in Equation 9:
SDMHR:= (yˆ if q′ ≥ q′min AND SDM(z′)yˆ ≥ ψyˆ⊥ otherwise).
Robustness and Out-of-Distribution Detection
Empirical results demonstrate that the SDM estimator is more robust to covariate-shifts and out-of-distribution inputs than existing classes of post-hoc calibration methods, while remaining informative over in-distribution data.
The value of q′min provides a principled, data-driven indicator of the reliability of the estimates,
serving as an interpretable check on whether class- and prediction-conditional estimates are obtainable at the specified α value. The SDMHR estimator is shown to be an effective outof-distribution detection method.
Updatability and Interpretability
The SDM activation function inherits the updatability property of instance-based metric learners,
meaning instances with labels y ∈ Y can be dynamically added to Dtr after training, which changes SIMILARITY and DISTANCE values without altering the MAGNITUDE or arg max prediction, provided the adaptor weights are fixed. This allows for a useful tradeoff between fast moving weights and slow moving weights in continual learning settings.
Furthermore, this approach provides interpretability-by-exemplar,
allowing users to examine applicable documents in training and calibration sets for further analysis.
Limitations
The SDM estimator requires the Dtr exemplar vectors at test-time, but the additional compute is comparable to existing dense retrieval mechanisms. The paper also notes that controlling for covariate shifts is more complex than label shifts and requires a partitioning of the high-dimensional space, which SDM estimators provide. However, it acknowledges that "it is not possible to maintain calibration over all possible distribution shifts.
Improvements for AI systems
As a fastidious researcher, I have analyzed the SIMILARITY-DISTANCEMAGNITUDE (SDM) activation function and its associated estimator, which is designed to quantify epistemic uncertainty in pre-trained language models (LMs).
Here are the specific improvements that can be implemented in AI systems using this framework:
-
Dominant Feature: Decomposing Epistemic Uncertainty into Three Interpretable Signals
-
Specific Improvement: Replace standard softmax activation with the SDM activation function, which explicitly decomposes predictive uncertainty into three distinct, actionable components for every prediction:
-
Specific Improvement 1 (SIMILARITY): Quantify the
Correctly predicted depth-matches into the training set.
This signal indicates how well a new instance aligns with known examples in the training data. -
Specific Improvement 2 (DISTANCE): Quantify the
DISTANCE to the training distribution,
normalized by class-wise empirical Cumulative Distribution Functions (CDFs) derived from a labeled calibration set. This measures how far an instance is from its nearest neighbors within its predicted class, providing a principled measure of distributional novelty. -
Specific Improvement 3 (MAGNITUDE): Quantify the
DISTANCE to the decision-boundary,
using the output of the exemplar adaptor's linear layer, which captures how close the prediction is to crossing a class threshold. -
System Capability: Enables
Interpretability-by-Exemplar via Dense Matching.
This allows researchers and users to identify exactly which training instances are most relevant (SIMILARITY) and how far an instance deviates from the learned manifold (DISTANCE), making model behavior transparent rather than opaque. -
Advanced Feature: Robust Selective Classification with the SDM Estimator
-
Specific Improvement: Implement the SDM estimator, which uses a data-driven partitioning of class-wise empirical CDFs to control class- and prediction-conditional accuracy among selective classifications.
-
System Capability: Provides
Robustness to Covariate Shifts and Out-of-Distribution Inputs.
Unlike existing calibration methods (like standard softmax or simple Platt scaling), the SDM estimator is empirically shown to maintain high conditional accuracy even when the input distribution shifts significantly, making it superior for real-world deployment where data drifts. -
Advanced Feature: High-Reliability Region Selection (SDMHR)
-
Specific Improvement: Use the derived metrics to define a
HIGH-RELIABILITY (SDMHR) region
via a principled selection criterion (Algorithm 1). This region is defined by simultaneously satisfying high similarity, low distance, and appropriate magnitude thresholds. -
System Capability: Facilitates
Triage in Decision Pipelines.
The system can automatically classify inputs into three categories:
12a. High-Reliability (SDMHR): Treat as automated/semiautomated decisions.
12b. Low-Reliability: Flag for human adjudication or resource-intensive LM tools.
12c. Rejected (R): Explicitly reject the prediction, indicating high epistemic uncertainty beyond the threshold of acceptable accuracy.
-
System Capability: Enhanced Out-of-Distribution (OOD) Detection
-
Specific Improvement: The SDM estimator is shown to reliably reject challenging OOD inputs in datasets like FACTCHECK and SENTIMENTOODSHUFFLED, where non-SDM methods perform poorly.
-
System Capability: Acts as an
Effective OOD Detection Method.
It provides a principled, data- and model-driven way to determine if a prediction is trustworthy based on its structural relationship to the training data manifold, rather than relying solely on task-specific heuristics.
In summary, this paper improves AI systems by transforming predictive output from a single probability score into a rich set of geometric and distributional metadata (Similarity, Distance, Magnitude), allowing for highly reliable selective classification and robust out-of-distribution detection in complex environments.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks