SCM-based Fairness and Faithful Explainability for Legal Document Classification

arXiv:2610.00045 · cs.CL · Submitted 2026-09-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "SCM-based Fairness and Faithful Explainability for Legal Document Classification".

Jane: SCM-based fairness regularisation for LegalBERT investigates whether debiasing interventions that change fairness also change how faithfully explanations reflect model reasoning.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to recap, the paper "SCM-based Fairness and Faithful Explainability for Legal Document Classification" sets out to test if applying an SCM regularization technique during the fine-tuning of LegalBERT can simultaneously adjust demographic fairness and the faithfulness of its explanations on a legal corpus.

Jane: That's right, Tom. It’s investigating that specific question: does changing how fair a model is also change how faithfully its explanations reflect what it actually did? They are testing this on the ECtHR alleged-violations corpus from LexGLUE.

Lu: The core setup involves comparing a standard LegalBERT baseline against an SCM-regularised version that specifically penalizes stereotypical warmth and competence representations during the fine-tuning process.

Meng: It seems like they're looking at three main areas: how well the model performs overall, whether it achieves demographic fairness, and how good its explanations are when we look at them closely.

Lalam: I think the most interesting part is their finding about a dissociation: that this debiasing intervention doesn't actually reduce demographic disparity but instead consistently degrades the quality of SHAP explanations without harming overall predictive performance.

The paper's summary: Tom: Exactly, Jane. The paper summarizes their findings by showing that while the SCM-based fairness regularisation didn't manage to lower demographic disparities on the ECtHR corpus, it did consistently degrade SHAP explanation sufficiency across five different random seeds.

Jane: That means they found that the intervention didn't reduce unfairness on either gender or ethnicity axes, but it still consistently made the explanations less sufficient when we checked them using perturbation-based sufficiency and comprehensiveness metrics.

Lu: The authors point out that a shuffled-pair control experiment, which used an arbitrary contrastive penalty instead of the SCM structure, reproduced all three patterns—flat performance, flat fairness, and degraded sufficiency—which helps pinpoint what's actually coming from the warmth–competence structure.

Meng: So they’re suggesting that whatever is happening here isn't just about the specific feature associations being targeted by SCM, but rather a general effect of contrastive representational regularization itself on the explanation quality.

Lalam: That points toward a really important implication: the resulting dissociation shows that fairness and explanation faithfulness are not coupled outcomes of a debiasing intervention, meaning we can’t just use explanation quality as a substitute for measuring fairness in this setting.

The paper's improvements: Tom: Now, looking at the proposed improvements, the authors suggest that because of this dissociation, the main takeaway is that explanation quality really cannot be used as a proxy for fairness when auditing legal-NLP systems.

Jane: They argue for establishing a direct fairness auditing protocol where we measure demographic disparities like DPD and EOD directly, rather than relying on metrics derived from SHAP explanations to infer whether bias mitigation actually worked.

Lu: Their suggestion is quite practical: the system should be audited against the "fairness null" demonstrated in their results; if DPD or EOD shows no significant improvement under SCM regularization, the audit report must state clearly that token-level stereotypical debiasing didn't shift disparity.

Meng: I see this as a strong call for better operational definitions of fairness because relying on explanation sufficiency seems to be a trap we need to avoid in high-stakes legal deployment scenarios.

Lalam: And they also point out a specific asymmetry: the degradation is consistent in sufficiency but comprehension remains unchanged, suggesting the intervention disrupts the top-token summary without affecting the full attribution set's contribution, which is a really nuanced observation.

Conclusion: Tom: So, wrapping up this discussion on "SCM-based Fairness and Faithful Explainability for Legal Document Classification," it seems like the main conclusion is that SCM regularisation doesn't reduce demographic disparity but consistently degrades explanation faithfulness without materially compromising predictive performance.

Jane: That really drives home the point about the dissociation they found, which establishes that in this legal-NLP setting, explanation faithfulness cannot serve as a proxy for fairness, demanding direct measurement of disparities.

Lu: I think it opens up a lot of creative avenues because since we know the structure is measurable through four quadrants recovering documented group stereotypes, we could potentially build more targeted debiasing methods operating directly on those learned representations.

Meng: From an engineering perspective, this means our next phase needs to focus on building those direct disparity measurement tools they advocate for, ensuring our monitoring systems focus squarely on the actual demographic differences rather than just explanation metrics.

Lalam: This work is important because it moves the needle away from using explanation quality as a proxy; it forces us to be explicit about what we are measuring when we claim a model is fair or trustworthy in legal contexts.

Tom: Fantastic summary, everyone. It sounds like we have a really clear direction now on how to approach fairness and explainability in these complex models. We’ve got some serious material here for the listeners who want to understand the real implications of this research on deploying AI responsibly in law and beyond.

Yasmina El Kacemi, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag

University of Amsterdam

cs.CL

Submitted: 2026-09-03

Updated: 2026-09-03

Importance score: 82/100

The gist: SCM-based fairness regularisation for LegalBERT investigates whether debiasing interventions that change fairness also change how faithfully explanations reflect model reasoning.

Key concepts

SCM Regularisation
A specific type of contrastive regularization added to the model's training loss. It forces the model to learn representations where antonym pairs (like 'law' and 'crime') are pushed apart in the embedding space, aiming to reduce biased associations.
Explanation Faithfulness
Measures how accurately a model's explanation (like SHAP values) reflects the actual reasoning behind a specific prediction. Sufficiency checks if only important tokens drive the prediction, and comprehensiveness checks if removing top tokens changes the result.
Demographic Disparity (DPD/EOD)
Metrics used to measure fairness by checking if model predictions are unequal across different demographic groups, such as gender or ethnicity. These metrics quantify whether certain groups receive positive classifications at different rates.
Dissociation
The central finding that a change in the explanation behavior (e.g., how tokens are weighted) did not correspond to a change in fairness metrics. This means the intervention targeted by SCM did not actually reduce demographic disparity.

Terminology

Summary

SCM-based fairness regularisation for LegalBERT investigates whether debiasing interventions that change fairness also change how faithfully explanations reflect model reasoning. This study compares a LegalBERT baseline against an SCM-regularised variant on the ECtHR alleged-violations corpus to determine if changes in explanation behaviour signal changes in fairness, finding a dissociation where SCM regularisation does not reduce demographic disparity but consistently degrades SHAP explanation sufficiency without materially compromising predictive performance.

The core investigation and its significance

This research investigates whether SCM-based fairness regularisation applied to LegalBERT affects (a) demographic disparity in predictions, (b) the faithfulness of SHAP explanations, and (c) classification performance on the ECtHR corpus from LexGLUE. The study is motivated by the unresolved question of whether a debiasing intervention that changes fairness also changes how faithfully explanations reflect the model’s reasoning. The contribution is a dissociation: a change in explanation behaviour did not signal a change in fairness, so fairness must be measured directly. This finding has direct implications for legal-NLP auditing, suggesting that explanation quality cannot serve as a proxy for fairness in this setting.

The experimental setup and models

The experiment employs four matched conditions: an unregularised LegalBERT baseline, an SCM-inspired contrastive regularised model, a shuffled-pair control, and a demographic word-pair debiasing control. All conditions are trained and evaluated under identical data, tokeniser, truncation, optimiser, and random seeds (42, 13, 7, 21, 100). The baseline is LegalBERT-Base-Uncased fine-tuned with a Binary Cross-Entropy loss weighted by log-smoothed inverse-frequency class weights. The SCM regularised model adds an SCM term to the classification loss: Ltotal = LCE + λ · LSCM, where LSCM is a contrastive margin loss over the antonym pairs present in each batch (Equation 3).

Evaluation metrics and fairness definitions

The dependent variables are predictive performance (macro F1), fairness (DPD and EOD, with DI diagnostic), and explanation faithfulness (perturbation-based sufficiency and comprehensiveness). Fairness is assessed using Demographic Parity Difference (DPD) measuring unequal positive-prediction rates, Disparate Impact (DI) as a diagnostic measure, and Equalized Odds Difference (EOD) which conditions on the ground-truth label. Explanation faithfulness is evaluated through sufficiency (magnitude of change in prediction probability when only top-k attributed tokens are retained), and comprehensiveness (change in prediction probability when top-k tokens are removed).

Key findings regarding performance, fairness, and faithfulness

The study reveals several key results:

  1. Classification performance is unchanged; the SCM model's macro F1 is 0.656 ± 0.012 compared to the baseline's 0.670 ± 0.013, with a mean difference of −0.014 across five seeds (Table 2).

  2. SCM regularisation does not reduce demographic disparity; mean DPD and mean EOD remain essentially unchanged on both gender and ethnicity axes under the SCM model compared to the baseline (Table 3). The null holds across two fairness definitions.

  3. The intervention consistently degrades SHAP explanation sufficiency across all five seeds, with a growth in sufficiency as k increases (from approximately +0.031 at k = 1% to +0.053 at k = 10%) (Table 4). This degradation is consistent under Integrated Gradients.

  4. A shuffled-pair control reproduces the three main patterns: flat performance, flat fairness, and degraded sufficiency, indicating the effect is attributable to contrastive representational regularisation of any kind, not the warmth and competence structure.

Conclusion on dissociation

The central finding is a five-seed dissociation: SCM regularisation leaves demographic disparity and classification performance largely unchanged while consistently degrading explanation faithfulness. This demonstrates that fairness and faithfulness are not coupled outcomes of a debiasing intervention, establishing that in this setting, explanation faithfulness could not have served as a proxy for fairness. The asymmetry is informative: sufficiency degrades while comprehensiveness remains unchanged, so the intervention disrupts the top-token summary without changing the full attribution set’s contribution. The results suggest that disparity is driven by factors beyond token-level stereotypical associations targeted by SCM.

Limitations and future directions

The conclusions are qualified by several limitations:

  1. A construct validity threat affects the fairness operationalisation, as protected groups are inferred from demographic keywords, which may not function as reliable identifiers in legal text.

  2. Faithfulness metrics carry a construct-validity limitation, as token-masking sufficiency assumes that model behaviour under masking generalises to natural input variation.

Improvements for AI systems

As a fastidious researcher, I have analyzed the findings of this study on SCM-based fairness regularisation for LegalBERT. The core contribution is the demonstration that debiasing interventions targeting latent stereotype structures (warmth/competence) do not reliably shift demographic disparities or improve classification performance, but they consistently degrade explanation faithfulness.

Based on these specific empirical results, here are the proposed improvements and the resulting capabilities of an enhanced AI system:


)

  1. Implement a Faithfulness-Aware Debiasing Layer in LegalBERT Fine-tuning:

  2. Develop a Post-hoc Explanation Confidence Monitor for High-Stakes Deployments:

  3. Establish a Direct Fairness Auditing Protocol (The Dissociation Principle):

  4. Implement a Faithfulness-Aware Debiasing Layer in LegalBERT Fine-tuning:

The AI system should be fine-tuned using the SCM contrastive regularisation term, but this must be coupled with a mechanism to monitor explanation quality during training.

  • When the SCM loss is minimized, simultaneously track the change in SHAP sufficiency and comprehensiveness (as shown in Table 4). If the regularization pushes explanations toward a state where sufficiency drops significantly (i.e., explanations become less faithful), an adaptive mechanism should dynamically adjust the regularization strength or introduce a secondary constraint to maintain a minimum acceptable faithfulness threshold.

  • This prevents the faithfulness degradation noted in Section 6 from becoming an unmonitored side effect, ensuring that any fairness intervention is not achieved at the expense of transparent reasoning.

  1. Develop a Post-hoc Explanation Confidence Monitor for High-Stakes Deployments:

The system should incorporate a real-time evaluation module that assesses the reliability of post-hoc explanations (SHAP/Integrated Gradients) before they are presented to legal practitioners or used in judicial support tools.

  • This monitor should calculate the Sufficiency Score (based on top-k token retention) and flag instances where this score falls below a pre-set threshold, indicating that the explanation is not reliably capturing the model's reasoning (H2 finding).

  • If a low sufficiency score is detected, the system should automatically trigger one of three actions:

List A: Reject the prediction and request human review.

List B: Re-run with a different attribution method (e.g., switch from SHAP to Integrated Gradients, as they showed directional consistency in Table 4).

List C: Flag the prediction as Low Faithfulness/High Uncertainty, providing a confidence score based on the explanation's reliability, rather than just the classification probability.

  1. Establish a Direct Fairness Auditing Protocol (The Dissociation Principle):

The system architecture must explicitly decouple fairness auditing from explanation fidelity, moving away from using explanation quality as a proxy for fairness.

  • Instead of relying on SHAP sufficiency metrics to infer bias mitigation, the protocol must mandate the direct measurement of demographic disparities (DPD and EOD) using robust operationalizations (like NER grouping for ethnicity).

  • The system should be audited against the Fairness Null demonstrated in Section 6: If DPD/EOD metrics show no statistically significant improvement under SCM regularization, the audit report must explicitly state that token-level stereotypical debiasing did not shift disparity.

  • This ensures that developers and legal users understand that a coherent explanation (high sufficiency) does not equate to an unbiased model, directly mitigating the risk of deploying a plausible but biased system.

The improved AI system will be:

  1. A legally accurate classifier with reduced reliance on potentially misleading token-level associations for bias mitigation.

  2. An inherently trustworthy system where explanation reliability is a primary quality metric, preventing the deployment of models whose reasoning cannot be faithfully represented by their explanations.

  3. A robust auditing tool that enforces a necessary separation between fairness and explainability, ensuring that fairness conclusions are based on direct disparity metrics rather than proxy metrics like SHAP sufficiency.

Sources

Related papers