Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

arXiv:2608.12935 · cs.AI, cs.LG · Submitted 2026-08-13 · Read on arXiv

Lei You

Technical University of Denmark

cs.AI, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/youlei202/decaf

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: DECAF: Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses Summary This paper introduces DECAF (Decomposition of Evidence, Contradiction, And Fragility), a method for

Terminology

Summary

DECAF: Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

Summary

This paper introduces DECAF (Decomposition of Evidence, Contradiction, And Fragility), a method for interpreting perturbation-based model explanations by decomposing the response magnitude into three semantically distinct components. The core problem identified is that perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual–counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint.

Formal Setup and Method

The method formalizes a paired information reveal where paired inputs (a factual x+ and counterfactual x−) are progressively revealed along matched trajectories from a common uninformative state x0. The signed response is r(t) = q(x+(t)) − q(x−(t)), with the final contrast d = r(1). DECAF uses the final contrast as a semantic reference: the gate a = 1 d ≥ ε determines whether the final effect is active, and the sign s = sign(d) provides orientation. The response is then routed as (e(t), c(t), f(t)) = (az+(t), az−(t), (1−a)r(t)), where z(t) = s·r(t). This produces three components: evidence E (endpoint-aligned responses), contradiction C (endpoint-opposed responses), and fragility F (responses on endpoint-null pairs). The decomposition is lossless: Abs = E + C + F exactly.

Theoretical Foundations

The paper proves three key theoretical results. Theorem 1 establishes that the routing is unique under three operational axioms: conservation (e + c + f = r), endpoint gating (f = 0 for active, e = c = 0 for null), and directional support (c = 0 when sr ≥ 0, e = 0 when sr ≤ 0). Theorem 2 shows that DECAF strictly refines response magnitude: "Ordinary response magnitude is a deterministic function of the DECAF profile, Abs = E + C + F, so every decision rule based on Abs can also be implemented from (E, C, F). The reverse recovery fails in general: for every m > 0, the distinct profiles (m, 0, 0), (0, m, 0), and (0, 0, m) all produce Abs = m." Proposition 1 demonstrates that contradiction separates attenuation from inversion: both can remove the same aligned evidence, but only inversion creates contradiction, with C/(E+C) = η in the inversion regime.

Controlled Behavioral Validation

The paper validates each component against independently measured behavior across controlled vision (3D Shapes) and tabular (Covertype) settings. For evidence, the evidence margin Ewall − Eshape correlates 0.936 with this reversal vulnerability across 52 training checkpoints, with 90% bootstrap interval [0.885, 0.961]. For fragility, the component tracks this prediction-change rate closely across robust, neutral, and fragile training variants. For contradiction, across 30 models and all mismatch levels, C tracks the pairwise label-swap rate with ρ = 0.961, whereas ordinary magnitude does not (ρ = −0.036). On Covertype with 135 classifiers, DECAF achieves Spearman correlations of 0.864 for preservation, 0.987 for actual inversion, and 0.974 for endpoint-null change, substantially outperforming baselines including native SHAP and SHAP interaction.

ImageNet-9 Real-World Audit

The paper conducts a 72-model audit on ImageNet-9 (24 off-the-shelf ImageNet-1k classifiers plus 48 models fine-tuned on four training variants). Key findings include:

  1. Behavioral alignment: Evidence achieves AUROC 0.930 for background reliance (vs 0.900 for magnitude), and fragility reaches AUROC 0.878 for endpoint-null sensitivity (vs 0.433 for magnitude and 0.316 for endpoint magnitude).

  2. Magnitude-matched comparison: Requiring their Abs values to differ by at most 5% yields 8,289 matched comparisons... Abs reaches a role-agreement accuracy of 0.350, whereas DECAF reaches 0.964. This demonstrates that nearly identical response magnitudes can correspond to different observed behaviors.

  3. Path dependence: Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4×. Specifically, patch reveal increases ordinary magnitude about 1.8×, with contradiction growing 1.8× and fragility more than 4×, while evidence remains close to its blend value. Model-rank transfer across paths shows evidence rankings remain stable (ρ = 0.86/0.77) while magnitude rankings are unstable (ρ = 0.17/0.26).

External Attribution Benchmarks and Scaling

On FunnyBirds and ImageNet-1k, DECAF trajectories outperform all tested general-purpose attribution baselines. DECAF-9 achieves 0.406 on FunnyBirds and 0.379 on ImageNet-1k, exceeding baselines like DeepLIFT (0.197/0.341), IG-32 (0.271/0.242), and RISE-512 (0.302/0.179). KernelSHAP reaches 0.447 on ImageNet-1k but is shown separately because it directly queries the same deletion intervention used to define the evaluation target. On the 1B-parameter DINOv2 ViT-g/14 model, DECAF-5 matches IG-32 quality (0.215 vs 0.213) while requiring 4.75× lower wall time and 2.36× lower peak memory.

The trajectory's value is clarified through ablations: on FunnyBirds (where evaluation uses different interventions than the explanation endpoint), the trajectory adds about 0.08 Spearman beyond endpoint-only M; on ImageNet-1k (where evaluation repeats the same deletion), five stages nearly match M while nine stages add only 0.007. The paper concludes: Magnitude tells us how much a model reacts; DECAF preserves that quantity while revealing what kind of response produced it.

Improvements for AI systems

Improvements to AI Systems:

  1. Interpretable Explanation Routing for Model Auditing
  • What it does: Replaces single scalar attribution scores with a three-channel output (Evidence, Contradiction, Fragility) for every input feature or perturbation step.

  • Capability: An AI auditor can now distinguish why a model changed its prediction—whether the change supports the final decision (evidence), opposes it (contradiction), or is an artifact of the perturbation path (fragility). This enables automated detection of models that rely on spurious correlations (high fragility) or that internally flip-flop (high contradiction) even when final accuracy is similar.

  1. Magnitude-Invariant Behavior Prediction
  • What it does: Uses the DECAF profile (E, C, F) instead of raw response magnitude (Abs) as input to downstream behavioral predictors.

  • Capability: An AI system can now predict model vulnerabilities (e.g., background reliance, label-swap sensitivity, endpoint-null instability) with far higher accuracy—e.g., AUROC 0.930 vs 0.900 for background reliance, and 0.878 vs 0.433 for endpoint-null sensitivity. It can also flag models with identical response magnitudes but opposite behaviors, which raw magnitude would miss entirely.

  1. Path-Robust Model Ranking
  • What it does: Computes DECAF evidence scores across multiple reveal paths (e.g., blend vs patch) and uses evidence, not magnitude, for ranking models.

  • Capability: An AI system can now rank models by robustness to input perturbations in a way that is stable across different perturbation strategies (rank correlation ρ=0.86/0.77 for evidence vs 0.17/0.26 for magnitude). This enables reliable model selection for deployment where perturbation type is unknown or varies.

  1. Contradiction-Aware Training Signal
  • What it does: Uses the contradiction component (C) as a regularizer or early-stopping metric during training.

  • Capability: An AI system can now detect and penalize models that internally reverse their decisions along the input-reveal path, even if final outputs are correct. This reduces hidden instabilities and improves generalization, especially in safety-critical applications where decision reversal is unacceptable.

  1. Efficient High-Quality Attribution with Trajectory Decomposition
  • What it does: Uses DECAF's trajectory-based decomposition (with 5–9 stages) as a drop-in replacement for expensive attribution methods like Integrated Gradients or SHAP.

  • Capability: An AI system can now generate attribution maps that match or exceed state-of-the-art quality (e.g., 0.406 on FunnyBirds, 0.379 on ImageNet-1k) while requiring 4.75× less wall time and 2.36× less peak memory than IG-32 on a 1B-parameter model. This enables real-time interpretability for large models in production.

  1. Endpoint-Null Sensitivity Detection
  • What it does: Uses the fragility component (F) to identify inputs where the model's response is entirely an artifact of the perturbation path and vanishes at the factual–counterfactual endpoint.

  • Capability: An AI system can now flag inputs where explanations are meaningless (fragility-dominant), preventing over-reliance on explanations for high-stakes decisions. This is particularly useful for detecting adversarial or out-of-distribution inputs where perturbation-based explanations are unreliable.

  1. Unified Behavioral Audit Framework
  • What it does: Integrates DECAF into a single pipeline that simultaneously outputs evidence, contradiction, and fragility scores for any model–dataset pair.

  • Capability: An AI system can now perform a comprehensive behavioral audit in one pass: detecting background reliance (via evidence), label-swap vulnerability (via contradiction), and path sensitivity (via fragility). This replaces multiple separate tests with a single, theoretically grounded, and empirically validated procedure.

Abstract

Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We introduce DECAF (Decomposition of Evidence, Contradiction, And Fragility), which routes aligned, opposed, and endpoint-null responses into evidence E, contradiction C, and fragility F. The decomposition preserves ordinary magnitude exactly, Abs = E + C + F, and is unique under endpoint-relative axioms. Across controlled vision and tabular settings, the three components track independently measured behavior. In a 72-model ImageNet-9 audit, we compare cases with nearly identical response magnitude but different independently measured behaviors. The largest DECAF component agrees with an observed behavior in 96.4% of cases, compared with 35.0% for magnitude alone. Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4x. On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform the tested general-purpose attribution baselines. On a 1B-scale DINOv2 model, a short trajectory matches a strong gradient-based baseline with 4.75x lower wall time and 2.36x lower peak memory.

Sources

Related papers