When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
Yang Liu, Ran Zou
University of North Carolina at Chapel Hill · University of California, Irvine
cs.CL, stat.ML
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: 16 pages, 1 figure
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
The gist: This paper studies black-box auditing of language model classifiers that return explanations alongside predictions.
Terminology
Summary
This paper studies black-box auditing of language model classifiers that return explanations alongside predictions. The authors investigate whether the relationship between a model's predicted label and its generated rationale becomes anomalous when a backdoor attack changes the predicted label, and whether this mismatch provides a lightweight signal for detecting backdoors.
The paper examines a deployment regime where defenders screen each incoming prediction individually, have only black-box query access and a small clean calibration set (n=256 examples), but do not have any information about the trigger. The defender can request a label together with a rationale or quoted evidence from the LM classifier. The attacker poisons the victim during fine-tuning so that inputs containing a hidden trigger are mapped to a target label, while clean-task accuracy remains high. Four OpenBackdoor-style attack families are evaluated: BadNets lexical triggers, AddSent trigger-sentence insertion, EP rare-token/embedding-poisoning triggers, and SOS stealth-oriented contextual triggers.
Groundedness Drift (Algorithm 1) is the default one-query audit. It queries the victim with a summary-rationale prompt to obtain a label and short answer summary, then measures whether the summary appears grounded in the input using a lightweight lexical score: DA(x) = 1 − max Glex, Gtfidf, where Glex is the fraction of non-stopword content tokens in the summary that also occur in the original input, and Gtfidf is the maximum TF-IDF cosine similarity between the summary and any clause of the input. Class-conditional thresholds are estimated from clean calibration data using victim-predicted label buckets.
Unsupported Groundedness (Algorithm 2) is an optional multi-probe escalation applied only after Groundedness Drift accepts an input. It issues a quoted-evidence prompt requesting exact input quotes, then runs evidence-only, evidence-keep, evidence-drop, and local document-view probes. It computes support-consistency drift (whether quoted evidence remains label-consistent under victim self-probes) and document-consistency drift (whether local views agree with the quoted-evidence label), converts components to class-conditional clean percentiles, and fuses them using the maximum.
Non-adaptive attacks: Across two 7B backbones (Qwen2.5-7B-Instruct and Mistral-7B-Instruct-v0.2), five datasets (SST-2, Rotten Tomatoes, TREC, Jigsaw, AG News), and four attack families, Groundedness Drift achieves higher AUROC and lower residual target ASR than every compared detector (ONION, BBCaL, CoS, and prompted-reasoning variants) in all cases at a nominal 5% clean-FPR budget. For example, on Qwen, Groundedness Drift reduces residual ASR to near zero on SST-2, Rotten Tomatoes, TREC, and Jigsaw, with realized clean FPR at most 3.7%. The weak CoS and Reasoning-best results show that generic rationale elicitation is insufficient; the useful signal is input grounding.
Adaptive explanation camouflage: When attackers jointly supervise the target label and a grounded-looking explanation on AG News, Groundedness Drift weakens significantly (AUROC 0.365–0.565). Unsupported Groundedness improves AUROC for every family and lowers residual ASR at both diagnostic budgets (10% and 20% clean FPR), but the gain is partial—residual ASR remains 0.462–0.633 at 10% clean FPR and 0.434–0.506 at 20%. The paper concludes that Unsupported Groundedness recovers useful signal without closing the adaptive gap.
The paper's central lesson is that explanations are useful traces, not certificates.
Auditing should distinguish whether an explanation is grounded in the input, whether that text operationally supports the label, and whether it actually caused the decision. The results find strong signal at the first level, partial signal at the second, and an open problem at the third. The authors recommend a layered audit approach: grounding screens routine traffic, while support and counterfactual probes examine higher-risk inputs. The open challenge is causally sensitive verification that preserves black-box access without requiring a costly trusted model.
Improvements for AI systems
Improvements to AI Systems Based on This Paper:
-
Add a lightweight
groundedness drift
pre-filter to any LLM classifier that returns explanations. The improved system computes a lexical overlap score (e.g., fraction of non-stopword tokens in the rationale that appear in the input, plus a TF-IDF clause similarity) for every prediction. If the score falls below a class-conditional threshold (calibrated on a small clean set), the system flags the input as suspicious and routes it to human review or a stronger detector—without needing to know the attack trigger. This reduces residual attack success rate to near zero for common backdoor families (lexical, sentence insertion, rare-token, contextual) at a 5% false-positive budget, while adding negligible latency (one extra query). -
Implement a two-stage audit pipeline: fast grounding screen, then targeted support-consistency probes for high-risk inputs. The improved system first applies the grounding check; only if it passes does it escalate to a multi-probe verification that requests exact quoted evidence from the input, then runs four self-probes (evidence-only, evidence-keep, evidence-drop, and local document-view) to measure whether the quoted evidence remains label-consistent. The system fuses these consistency scores into a percentile-based anomaly score. This recovers partial detection signal even against adaptive attackers who craft grounded-looking explanations, reducing residual attack success rate from 0.9 to 0.5 at a 10% false-positive budget—useful for triage in high-stakes domains (e.g., content moderation, medical triage, legal document review).
-
Build a
rationale grounding
metric into model evaluation and monitoring dashboards. The improved system continuously tracks the distribution of grounding scores (lexical and TF-IDF) per predicted label over time. Sudden shifts in this distribution—even without known attacks—trigger alerts for data drift, prompt injection, or poisoning. This provides a cheap, model-agnostic health signal that complements accuracy-based monitoring, especially in black-box deployment where internal weights are inaccessible. -
Enhance adversarial robustness by training a separate
grounding verifier
model. The improved system trains a small, fast classifier (e.g., a fine-tuned RoBERTa) to predict whether a rationale is genuinely grounded in the input, using synthetic negative examples (rationales with swapped, paraphrased, or unrelated content). This verifier can be deployed as a real-time gate before any LLM explanation is shown to end users, reducing the risk of misleading or hallucinated justifications in production. -
Design a
counterfactual probe
module for causal verification of explanations. The improved system, after detecting a grounded but potentially unsupported explanation, automatically generates minimal input edits (e.g., removing the quoted evidence clause) and re-queries the victim to see if the label flips. If the label does not flip, the explanation is flagged as non-causal. This addresses the paper's open challenge at the third level (causal sensitivity) and can be implemented with a small number of additional queries, enabling black-box causal audits without a trusted model.
What the improved AI system can do:
-
Detect backdoored LLM classifiers in real time with high accuracy and low false positives, using only black-box queries and a small clean calibration set.
-
Triage high-risk inputs for deeper inspection, catching even adaptive attacks that generate plausible explanations.
-
Provide continuous, explanation-aware monitoring for drift or poisoning in deployed models.
-
Prevent users from being misled by unsupported or non-causal rationales, improving trust and safety in LLM-based decision support.
-
Operate without access to model weights or triggers, making it practical for API-based and third-party model deployments.
Sources
- The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Mistral 7B
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering