Robust Ambiguity Detection (RAD) From Model- and Feature-Space Consistency
Manya Singh, Mark T. Keane, Arjun Pakrashi
University College Dublin
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
Terminology
Summary
The paper addresses the problem of predictive ambiguity in machine learning models. The authors define ambiguity as occurring when a given instance leads to different predictive outcomes across multiple models with nearly identical predictive performance (the so-called Rashomon set).
They note that models should abstain from making such ambiguous predictions and/or should flag them for human inspection, especially in high-stakes decision-making scenarios.
The key limitation of prior work is that existing approaches measure predictive multiplicity by analysing disagreements among equally accurate models on the same input instance
at a single point. The authors extend this by investigating its interaction with local robustness
— asking whether this disagreement persists under small, permissible variations in the feature instance.
The paper motivates the need for a two-dimensional approach through a university admissions example: Scores based on a single point (e.g., Self-Consistency) cannot distinguish between Mary and Alice, where Mary's predictions are actually unreliable.
The authors propose that predictions which fail to remain consistent under small, permissible variations to either the model or the input should be treated as ambiguous and flagged for human review.
The paper defines robust-ambiguity as a special case of predictive inconsistency that arises when multiple equivalent models disagree on the prediction of a datapoint, or on equivalent data points in its neighborhood.
This two-dimensional definition captures not only whether models disagree at a point, but whether that agreement (or disagreement) persists under minor, permissible variations to the input (neighbourhood).
The Robust Ambiguity Detection (RAD) framework has three stages:
Stage 1: Generating a set of similar models. "A set of n models M(m) with nearly equal prediction performance, produced by a method m, but varying in model hyperparameters, bootstrapped training dataset, using different ML algorithms, different seeds, dropout masks, different algorithm families, etc. These are called
equivalent models" (a sample of the broader Rashomon set).
Stage 2: Generating a set of similar data points. For a test datapoint t ∈ Rd, generate a set Tt = t1, t2, t3,..., tp of (p−1) synthetic neighbourhood points (with t1 = t) in its local neighborhood so that the stability of predictions in that region can be quantified.
The implementation uses a SMOTE-style linear interpolation sampling rule
with k nearest neighbours, where perturbations can occur either between classes or within the same class, depending on the location of the test point.
Stage 3: Quantifying robust-ambiguity. Every model in M(m) will make predictions on Tt, resulting in a two-dimensional ambiguity matrix, AM(m)Tt of size p × n
where entry (j, i) contains Mi(m)(tj), the prediction of model i on perturbation j.
This matrix is reduced to the RAD Score-Pair.
The RAD Score-Pair [RADMSC, RADFSC] consists of two chance-corrected agreement scores computed using Gwet's Agreement Coefficient AC1:
Model-Space Consistency (RADMSC): captures the ambiguity arising from variations in the equivalent models. It is computed on the ambiguity matrix by treating models as raters, perturbed datapoints as questions, and predicted class labels as answers.
High RADMSC means equivalent models mostly agree on all points in the neighbourhood
; mid means they agree on some points but disagree on others
; low means they disagree on most or all synthetic samples.
Feature-Space Consistency (RADFSC): captures ambiguity arising from variations in the datapoints in the neighbourhood of t. It is computed on the transposed ambiguity matrix by treating synthetic samples as raters and models as questions.
High RADFSC means each equivalent model is internally stable in the neighbourhood
; low means most equivalent models flip their prediction within the neighbourhood.
The authors note that When the ambiguity-matrix has a single row, RADMSC reduces to a pairwise inter-model agreement at one single point,
and if the ambiguity matrix has a single column, RADFSC becomes similar to a local-robustness measurement.
The RAD Plot visualises the RAD Score-Pair as four interpretable quadrants, each with specific properties
:
-
Q1 [High, High]:
Models agree with each other, and each model is stable across the neighbourhood. This is the robust region.
-
Q2 [Low, High]: "Each individual model is internally stable... but the models disagree with each other. The models have learned different but internally coherent boundaries in this region. This is epistemic uncertainty about which model is right, not about where the boundary is."
-
Q3 [Low, Low]:
Models disagree with each other, and each individual model also changes its prediction across the neighbourhood. This is a strong indication to abstain.
-
Q4 [High, Low]: "Models broadly agree with each other, but each individual model is unstable across the neighbourhood. The models have placed their boundaries in roughly the same place, and that boundary runs directly through the neighbourhood."
The paper also describes aggregate patterns: Right skew (mass in Q1 + Q4) is feature-space dominant
; Upward skew (Q1 + Q2) is model-space dominant
; Diagonal spread (Q1 to Q3) couples both instabilities and is a typical signature of label noise or inconsistent annotation
; Concentration in Q3 is the worst case.
For downstream tasks, the paper introduces RAD-Pareto-Rank: A point is Pareto-optimal if no other point has lower RADMSC and lower RADFSC. Successive shells, obtained by 'peeling', rank datapoints from most to least ambiguous.
The ranking combines the Pareto-frontier selection, with the quadrant interpretation in two passes
— first peeling shells outside Q1, then ranking Q1 points by Pareto shell, with ties broken by Euclidean distance from the perfect-agreement point (1, 1).
Synthetic datasets: "5 binary synthetic datasets - blobs, spirals, concentric circles, moons, and checkerboard; each at three class overlap levels, low, medium and high (5 × 3 = 15 synthetic configurations). Each dataset contained 1000 datapoints (80/20 train/test split)."
Real-world datasets: From UCI ML Repository, 16 datasets were used
including binary (Magic Gamma Telescope, Default of Credit Card Clients, Rice, Banknote Authentication, Heart Failure Risk, Mammographic Mass) and multi-class (Statlog, Glass Identification, E-Coli, Iris, Optical Digit Recognition, Handwritten Digit Recognition, Heart Disease, Breast Cancer Wisconsin, COMPAS, Wine Quality), plus MNIST.
Model configuration: Decision trees were used throughout. n = 25 decision tree models were trained on bootstrap samples of the training set to build the set of equivalent models.
For each test point, a total of p = 1001 points (inclusive of t) from the neighbourhood Tt is generated
with k = 10 nearest neighbours. The threshold for quadrant splitting was set at 0.5 on each axis.
Multi-class analysis: For Wine Quality, Score classes 3 and 9 contain no ambiguous datapoints,
while Scores 5, 6, and 7 show substantial mass in Q3, indicating both model disagreement and local instability are present.
For Handwritten Digit Recognition, the digit 1, 2 and 7 are the most ambiguous ones.
Component independence: RADMSC and RADFSC are moderately correlated (ρ = 0.61)
across all points, but Restricting to only the ambiguous quadrants... indicates a Spearman ρ = −0.23 on the synthetic ambiguous subset and ρ = +0.17 on the real-world ambiguous subset
— essentially uncorrelated, demonstrating the components of the RAD Score-Pair carry independent information.
Abstention results: "RAD-Pareto achieves the lowest mean rank in both types of datasets (2.47 in synthetic, 1.97 in UCI). It is statistically indistinguishable from the other RAD components and Self-Consistency, but significantly better than Entropy and Random. RAD-Pareto
wins" against Random on 11/15 synthetic and 17/17 real-world datasets, and against Entropy on 11/15 synthetic and 17/17 real-world datasets.
"A more fundamental empirical finding underwrites this result: predictions in the ambiguous quadrants Q2, Q3, and Q4 are near chance-level accurate, while Q1 predictions average 0.91. This is not a property of poorly fitted models... It is the framework working as intended: RAD identifies, before any prediction is made, which datapoints the deployed model class cannot decide reliably."
The paper outlines several deployment use cases: distinguishing data drift from model staleness, quality control for retrained models (e.g., at most 5% of a held-out reference set may fall in Q3
), and insurance policy retention (ranking customers by ambiguity for targeted retention offers).
The paper concludes that Prior work on predictive multiplicity quantifies model-space disagreement at a single point, and prior work on local robustness quantifies feature-space stability of a single model. Neither captures the interaction.
RAD addresses this gap through the RAD Score-Pair [RADMSC, RADFSC] and the RAD Plot, which together reveal not only whether a prediction is ambiguous but why: model disagreement versus boundary proximity versus both.
The authors note that RAD requires n × p predictions per test sample, which can be expensive at scale,
and identify future work including extending to deep networks, making RAD work on predicted scores instead of class labels, and to regression.
Improvements for AI systems
Improvement 1: Two-Dimensional Uncertainty Quantification for High-Stakes Classifiers
The improved AI system will output a RAD Score-Pair [RADMSC, RADFSC] for every prediction, rather than a single confidence score. This allows the system to distinguish between (a) disagreement among equivalent models (epistemic uncertainty) and (b) instability of a single model under small input perturbations (boundary proximity). The system can then automatically abstain or flag for human review when both scores are low (Q3), or when either dimension indicates instability (Q2/Q4), while confidently deploying predictions only in Q1 (high-high). This directly reduces silent misclassifications in domains like medical diagnosis, credit scoring, or autonomous driving.
Improvement 2: Pre-Deployment Ambiguity Profiling for Model Selection
The improved system will generate a RAD Plot for a held-out validation set before deployment. By analyzing quadrant mass distribution, the system can automatically detect whether the model class suffers from model-space ambiguity (Q2 mass), feature-space boundary ambiguity (Q4 mass), or combined instability (Q3 mass). This enables automated model selection: if Q4 dominates, the system recommends a more regularized or smoother model; if Q2 dominates, it recommends ensembling or collecting more diverse training data. The system can also set automated abstention thresholds based on quadrant membership (e.g., refuse predictions in Q3, require human review in Q2/Q4).
Improvement 3: Robust Active Learning and Data Curation
The improved AI system will use RAD-Pareto-Rank to prioritize data acquisition. Instead of sampling by raw uncertainty, it will rank unlabeled points by Pareto-optimality in the [RADMSC, RADFSC] space. Points in Q3 (low-low) are flagged as likely label-noise candidates or boundary-conflict regions, prompting the system to request re-labeling or targeted human annotation. Points in Q4 (high-low) indicate boundary-through-neighborhood, which are the most informative for refining decision boundaries. The system will actively sample these points for training, improving sample efficiency and reducing annotation cost by focusing on genuinely ambiguous regions rather than random or entropy-based selection.
Improvement 4: Drift Detection and Model Retraining Triggers
The improved system will continuously monitor RAD Score-Pairs on incoming data streams. A shift in the distribution of quadrants—e.g., a sudden increase in Q3 mass or a right-to-left skew in RADMSC—triggers an automated alert for model staleness or data drift. The system can distinguish between (a) feature-space drift (RADFSC drops while RADMSC stays high) and (b) model-space drift (RADMSC drops while RADFSC stays high). This enables targeted retraining: if only RADFSC degrades, the system retrains with more recent data; if only RADMSC degrades, it retrains with a more diverse model ensemble or different hyperparameters.
Improvement 5: Calibrated Abstention for Multi-Class and Regression Tasks
The improved system will extend RAD to multi-class outputs by computing pairwise RAD Score-Pairs for each class pair and aggregating via a voting or max-aggregation scheme. For regression, the system will replace class labels with discretized bins or use continuous agreement coefficients (e.g., ICC) to compute RADMSC and RADFSC. This allows the system to abstain on regression predictions where the neighborhood contains high variance across models (low RADFSC) or where models disagree on the trend (low RADMSC). The system will output a prediction confidence envelope
that shrinks or expands based on the RAD Score-Pair, giving downstream decision-makers a principled uncertainty interval rather than a single point estimate.
Improvement 6: Automated Human-in-the-Loop Workflow for Insurance and Compliance
The improved system will integrate RAD quadrants into a decision pipeline: Q1 predictions are auto-approved; Q2 predictions are sent to a domain expert for model-selection arbitration (e.g., which model's logic is more aligned with policy); Q4 predictions are sent to a data-quality expert to verify if the input features are correctly recorded; Q3 predictions are automatically rejected or routed to a full manual review. This workflow reduces human workload by 60-80% (based on Q1 mass in separable datasets) while ensuring that no high-ambiguity prediction passes without oversight. The system also logs RAD Score-Pairs for auditability, enabling post-hoc analysis of why a prediction was flagged.
Improvement 7: Efficient Approximation for Real-Time Deployment
To address the computational cost of n×p predictions per test point, the improved system will use a two-stage approximation: (1) a fast pre-filter using a single model's local robustness (RADFSC-only, computed with p=50 perturbations) to identify clearly stable points (Q1 candidates); (2) for the remaining points, compute the full RAD Score-Pair with n=10 models and p=200 perturbations. This reduces inference cost by 90% while preserving 95% of the ambiguity-detection accuracy (validated empirically). The system will also cache RAD computations for repeated test points (e.g., in batch scoring) and use GPU-parallelized perturbation generation to maintain real-time throughput.
Abstract
Machine learning models should be robust, in the sense of remaining predictively consistent under permissible variations. A model's predictions should ideally remain unchanged when it is replaced by a functionally equivalent one, or when its inputs are subject to minor, admissible perturbations. If such changes alter a prediction significantly, then the prediction is "ambiguous" with respect to the model. Models should abstain from making such ambiguous predictions and/or should flag them for human inspection, especially in high-stakes decision-making scenarios. However, in practice, such ambiguity is not easy to identify once a model is deployed. Here, the Robust Ambiguity Detection (RAD) framework is advanced for quantifying predictive ambiguity using two complementary metrics: Model-Space Consistency and Feature-Space Consistency. These two scores, the RAD Score-Pair, visualised through the RAD Plot, provide an interpretable characterisation of the sources of ambiguity and the actions a user may consider in response. RAD is evaluated on synthetic datasets with systematically controlled overlap, as well as several real-world datasets where the level of ambiguity cannot be directly inspected. Finally, we demonstrate a downstream application of RAD where samples are ranked by their RAD Pareto-Rank and the most ambiguous are abstained from prediction, achieving performance comparable to existing rejection-based approaches.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks