Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information

arXiv:2608.10766 · cs.AI, cs.LG, stat.ML · Submitted 2026-08-13 · Read on arXiv

Kaivalya Rawal, Daria Onitiu, Brent Mittelstadt, Sandra Wachter, Chris Russell

University of Oxford · Hasso Plattner Institute · Weizenbaum Institute

cs.AI, cs.LG, stat.ML

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Code available at: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

Code: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

Project page: https://kairawal.github.io/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: The paper proposes "Rule of Thumb" (ROT) explanations, a new approach to Explainable Artificial Intelligence (XAI) that "identifies the most relevant features for predicting the behaviour of an AI

Terminology

Summary

The paper proposes Rule of Thumb (ROT) explanations, a new approach to Explainable Artificial Intelligence (XAI) that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. The authors state: "We propose 'Rule-of-Thumb' (ROT), a novel form of XAI that considers important inputs to be those that are most predictive of the outputs of an AI system. Rather than formulating XAI in terms of how altering an input value would alter the system output, we ask how knowledge of particular feature values should alter our predictions of the AI system's behaviour."

The paper presents the following mathematical formulation. Given a model C(·) defined over features J, and dataset X, ROT finds the optimal set of additive functions such that for any subset of features J ⊆ J and datapoint x ∈ X we estimate C(x) using their sum:

C(x) ≈ F(∑ fθj(xj) + G) ∀J ∈ P(J)

where "F is a sigmoid function (for classification) or an identity function (for regression), fθj(xj) is a learnt function representing the information provided by knowing that the jth feature takes value xj, and G a global bias term that indicates what prediction should be made without knowing anything about the datapoint x."

The optimization objective is:

min L = min ∑ x∈X ∑ J⊆J p J(1-p)(J-J) · l[C(x), F(∑ j∈J fθj(xj) + G)]

The authors note that Although computing L exactly is expensive and requires computing the standard loss over the entire dataset 2 J times, it can be efficiently optimized using dropout applied to the feature importances fθj(xj).

The paper identifies three critical assumptions of sensitivity-based explainers (like SHAP and LIME) that ROT avoids: "(i) that it is possible to query the AI system with new datapoints; (ii) the system responds similarly to synthetically perturbed data in the same way it does to real-world data, and (iii) The outputs of a system are continuous, and typically vary when a limited number of features are changed. ROT makes none of these assumptions."

The authors explain: "Sensitivity-based approaches LIME and SHAP optimise a per-datapoint objective. While ROT masks information and fits one model for the entire dataset, these approaches perturb datapoints and fit one linear model per datapoint."

The paper demonstrates ROT's utility in three zero-shot classification scenarios:

Judicial Case Predictions: Using a fine-tuned RoBERTa model for Indian court appeal outcomes, ROT achieved a weighted AUROC of 0.77 against human annotations, compared to SHAP (0.74), Integrated Gradients (0.65), LIME (0.62), and random (0.47).

Movie Review Classification: Using OpenAI's GPT-4.1-nano API (accessible only via API), ROT achieves a weighted AUROC of 0.72, while a random baseline only achieves 0.50. The authors note that Gradient based explainers (Integrated Gradients) cannot be compared, as access to GPT-4.1-nano weights is restricted, and perturbation based explainers (LIME, SHAP) are infeasible to apply even at this modest scale.

Resume Filtering: Using GPT-4.1-nano API, ROT explained predictions in terms of both input features and non-input features (race, gender, political orientation). The authors verify findings from the original study indicating that, on this dataset, GPT ascribes negligible importance to race, gender, and political orientation when making hiring decisions.

The paper replicates The Markup's audit of Amazon's e-commerce recommendation system. The authors demonstrate that explanations from this commonly employed workaround can be contested by varying the mimic model used; whereas ROT directly provides explanations without depending on mimic models.

They found that different equally-performant mimic models (all with test accuracy between 0.69-0.72) produced different SHAP and LIME explanations. The authors state: This 'different local behaviour but equivalent performance' casts doubt on the mimic approach. An audit using a particular mimic can be undermined by adversarially selecting a different mimic.

Critically, "ROT also uncovers 'sold by amazon' to be an important feature, something that remained hidden in preceding mimic based analysis. This is a new insight signalling potentially self-preferencing behaviour where the Amazon platform does not just promote its own brands, but also third party product sold directly by Amazon over the same product sold by third-party merchants."

The paper notes: Upon analyzing 1151 research articles at the intersection of XAI and science, we found 4.9% of them used explanations of AI systems as a means of proposing novel scientific hypotheses. Of these, 70% used SHAP.

Identifying Uninformative Features: Using a diabetes prediction dataset transformed via FairPCA so that age is uncorrelated with diabetes, ROT is the only explainer to consider age unimportant. SHAP and LIME both lead to spurious hypotheses that indicate the importance of age in diabetes prediction.

Adversarial Robustness: The paper replicates adversarial attacks from Slack et al. that hide sensitive features behind foil features. Results show: "Across 10 experiments, for LIME or SHAP the rate of recovering the sensitive feature as most important does not exceed 5% except once... ROT has a success rate always greater than 89%, and in 6 of the 10 experiments ROT shows a perfect 100% recovery rate."

The paper reports dramatic computational advantages: "generating the first ROT explanation took 1160 seconds, and each additional explanation took less than 0.1 milliseconds. By contrast, the average time taken for each SHAP explanation was 13 million times longer at 1320 seconds."

The authors explain: "Both SHAP and LIME need to fit a new interpretable model for each prediction they explain. This requires data collection by making arbitrarily time-consuming calls to the AI system for each additional explanation. By contrast, ROT uses dropout to fit a single interpretable ensemble on the entire unperturbed dataset, and each explanation is generated using per-feature importance scores from this model."

The paper argues ROT meets specific requirements from leading AI regulations. Specifically:

  • "ROT's narrow focus on directly predicting the behaviour of an existing system on real data, gives it robustness against adversarial attacks, and removes the need for mimic models. This gives it a clear advantage in meeting the EU AI Act's requirements for accuracy, robustness, and adversarial testing of high-risk AI systems under Article 15."

  • "ROT can be computed without access to the original AI system or the ability to obtain outputs for new inputs. This makes ROT valuable for enforcing regulations which do not grant access to third-party AI systems or source code. For example under Article 74(12) of the EU AI Act, regulators are granted access to observe AI outputs on fixed input data, but not the ability to obtain new predictions on arbitrary inputs."

  • "ROT can help identify self-preferencing practices, such as search result manipulation, that may be illegal under horizontal competition laws and the EU Digital Markets Act (DMA) that require establishing a link between model behaviour and anti-competitive practices."

The paper concludes: "ROT addresses several long-standing concerns in the community about the validity and practical utility of XAI, and importantly does so without making additional assumptions. It is flexible and useful in a range of emerging XAI applications areas where other forms of XAI simply can not be used – evaluating LLMs and explaining their behaviour in zero-shot classification, auditing proprietary AI systems where we can only observe their predictions and not alter input features, and using XAI to understand real-world phenomena and propose novel scientific hypotheses."

The code is available at: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

Improvements for AI systems

Based on this paper, here are the specific improvements I can make to AI systems:

  • Improvement: I can implement ROT as a post-hoc explainer that requires only logged input-output pairs, not live query access or gradient information.

  • Capability: The improved system can explain any AI model (including proprietary LLMs like GPT-4.1-nano, Amazon recommendation engines, or closed-source classifiers) using only historical data, without needing to perturb inputs or access internal weights.

  • Improvement: I can replace sensitivity-based attribution (SHAP/LIME) with ROT's information-based masking approach, which is resistant to feature-hiding attacks.

  • Capability: The improved system can recover truly predictive features even when sensitive attributes are deliberately obscured by correlated foil features—achieving >89% recovery vs. <5% for SHAP/LIME in adversarial settings.

  • Improvement: I can train a single ROT model per dataset using dropout-based optimization, then generate explanations for any number of datapoints in microseconds.

  • Capability: The improved system can explain millions of predictions in real-time (e.g., for fraud detection, medical triage) where per-instance SHAP/LIME would be computationally prohibitive (13 million× slower per explanation).

  • Improvement: I can build an audit tool that uses ROT to explain a proprietary system's behavior using only fixed, pre-collected inputs (as allowed under EU AI Act Article 74(12)).

  • Capability: The improved system can detect self-preferencing (e.g., Amazon promoting sold by Amazon products) or discriminatory hiring patterns without needing to query the target system with new inputs—enabling compliance checks on systems where regulators lack API access.

  • Improvement: I can apply ROT's global, dataset-level explanation to identify features that genuinely predict model behavior, rather than local perturbations that confound with correlated variables.

  • Capability: The improved system can propose valid scientific hypotheses (e.g., in diabetes research) by correctly ignoring age when it's uncorrelated with the outcome—where SHAP/LIME would falsely flag it as important, leading to wasted research effort.

  • Improvement: I can use ROT to explain LLM decisions in zero-shot classification tasks (legal judgments, movie reviews, resume screening) using only API outputs.

  • Capability: The improved system can audit LLM bias (e.g., detecting if race/gender/political orientation are actually used in hiring decisions) and provide human-interpretable explanations that outperform random baselines (AUROC 0.72 vs 0.50) without needing gradient access.

  • Improvement: I can extend ROT to handle both input features and non-input features (e.g., demographic attributes not present in the data) in the same explanation framework.

  • Capability: The improved system can answer questions like Does this resume model use gender? even when gender isn't a model input—enabling fairness audits that current explainers cannot perform.

  • Improvement: I can use ROT to provide consistent, non-contestable explanations that don't vary with the choice of mimic model.

  • Capability: The improved system can produce auditable explanations that withstand adversarial challenges (e.g., a company cannot dispute an audit by selecting a different equally-performant mimic model), making AI governance more reliable.


Bottom line: The improved AI system can explain any black-box model using only passive observation, resist adversarial manipulation, scale to millions of explanations, meet regulatory audit requirements, and generate scientifically valid hypotheses—all without the assumptions (query access, perturbation validity, output continuity) that break current explainers.

Abstract

Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. We show how RoT is well-suited to enable XAI in: (a) zero-shot classification using large language models (LLMs), (b) auditing of opaque AI systems without model access, and (c) the use of AI in scientific discovery. Additionally, RoT meets specific requirements from leading AI regulations, provides a familiar interface and visualisations for XAI practitioners, is model-agnostic, and is substantially faster than alternatives. Code available at: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

Sources

Related papers