SCM-based Fairness and Faithful Explainability for Legal Document Classification

summary

Video file (mp4)

The gist

SCM-based fairness regularisation for LegalBERT investigates whether debiasing interventions that change fairness also change how faithfully explanations reflect model reasoning.

In short

Researchers tested if using SCM-based fairness regularization in LegalBERT changed how well explanations reflected model reasoning. The study found that while fairness metrics like demographic disparity remained unchanged, the method consistently degraded explanation faithfulness, showing that explanation quality is not a reliable proxy for fairness.

Key concepts

SCM Regularisation
A specific type of contrastive regularization added to the model's training loss. It forces the model to learn representations where antonym pairs (like 'law' and 'crime') are pushed apart in the embedding space, aiming to reduce biased associations.
Explanation Faithfulness
Measures how accurately a model's explanation (like SHAP values) reflects the actual reasoning behind a specific prediction. Sufficiency checks if only important tokens drive the prediction, and comprehensiveness checks if removing top tokens changes the result.
Demographic Disparity (DPD/EOD)
Metrics used to measure fairness by checking if model predictions are unequal across different demographic groups, such as gender or ethnicity. These metrics quantify whether certain groups receive positive classifications at different rates.
Dissociation
The central finding that a change in the explanation behavior (e.g., how tokens are weighted) did not correspond to a change in fairness metrics. This means the intervention targeted by SCM did not actually reduce demographic disparity.

Terminology used across episodes

This episode discusses

The paper

SCM-based Fairness and Faithful Explainability for Legal Document Classification · Read on arXiv

Yasmina El Kacemi, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag

University of Amsterdam

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "SCM-based Fairness and Faithful Explainability for Legal Document Classification".

Jane: SCM-based fairness regularisation for LegalBERT investigates whether debiasing interventions that change fairness also change how faithfully explanations reflect model reasoning.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to recap, the paper "SCM-based Fairness and Faithful Explainability for Legal Document Classification" sets out to test if applying an SCM regularization technique during the fine-tuning of LegalBERT can simultaneously adjust demographic fairness and the faithfulness of its explanations on a legal corpus.

Jane: That's right, Tom. It’s investigating that specific question: does changing how fair a model is also change how faithfully its explanations reflect what it actually did? They are testing this on the ECtHR alleged-violations corpus from LexGLUE.

Lu: The core setup involves comparing a standard LegalBERT baseline against an SCM-regularised version that specifically penalizes stereotypical warmth and competence representations during the fine-tuning process.

Meng: It seems like they're looking at three main areas: how well the model performs overall, whether it achieves demographic fairness, and how good its explanations are when we look at them closely.

Lalam: I think the most interesting part is their finding about a dissociation: that this debiasing intervention doesn't actually reduce demographic disparity but instead consistently degrades the quality of SHAP explanations without harming overall predictive performance.

The paper's summary: Tom: Exactly, Jane. The paper summarizes their findings by showing that while the SCM-based fairness regularisation didn't manage to lower demographic disparities on the ECtHR corpus, it did consistently degrade SHAP explanation sufficiency across five different random seeds.

Jane: That means they found that the intervention didn't reduce unfairness on either gender or ethnicity axes, but it still consistently made the explanations less sufficient when we checked them using perturbation-based sufficiency and comprehensiveness metrics.

Lu: The authors point out that a shuffled-pair control experiment, which used an arbitrary contrastive penalty instead of the SCM structure, reproduced all three patterns—flat performance, flat fairness, and degraded sufficiency—which helps pinpoint what's actually coming from the warmth–competence structure.

Meng: So they’re suggesting that whatever is happening here isn't just about the specific feature associations being targeted by SCM, but rather a general effect of contrastive representational regularization itself on the explanation quality.

Lalam: That points toward a really important implication: the resulting dissociation shows that fairness and explanation faithfulness are not coupled outcomes of a debiasing intervention, meaning we can’t just use explanation quality as a substitute for measuring fairness in this setting.

The paper's improvements: Tom: Now, looking at the proposed improvements, the authors suggest that because of this dissociation, the main takeaway is that explanation quality really cannot be used as a proxy for fairness when auditing legal-NLP systems.

Jane: They argue for establishing a direct fairness auditing protocol where we measure demographic disparities like DPD and EOD directly, rather than relying on metrics derived from SHAP explanations to infer whether bias mitigation actually worked.

Lu: Their suggestion is quite practical: the system should be audited against the "fairness null" demonstrated in their results; if DPD or EOD shows no significant improvement under SCM regularization, the audit report must state clearly that token-level stereotypical debiasing didn't shift disparity.

Meng: I see this as a strong call for better operational definitions of fairness because relying on explanation sufficiency seems to be a trap we need to avoid in high-stakes legal deployment scenarios.

Lalam: And they also point out a specific asymmetry: the degradation is consistent in sufficiency but comprehension remains unchanged, suggesting the intervention disrupts the top-token summary without affecting the full attribution set's contribution, which is a really nuanced observation.

Conclusion: Tom: So, wrapping up this discussion on "SCM-based Fairness and Faithful Explainability for Legal Document Classification," it seems like the main conclusion is that SCM regularisation doesn't reduce demographic disparity but consistently degrades explanation faithfulness without materially compromising predictive performance.

Jane: That really drives home the point about the dissociation they found, which establishes that in this legal-NLP setting, explanation faithfulness cannot serve as a proxy for fairness, demanding direct measurement of disparities.

Lu: I think it opens up a lot of creative avenues because since we know the structure is measurable through four quadrants recovering documented group stereotypes, we could potentially build more targeted debiasing methods operating directly on those learned representations.

Meng: From an engineering perspective, this means our next phase needs to focus on building those direct disparity measurement tools they advocate for, ensuring our monitoring systems focus squarely on the actual demographic differences rather than just explanation metrics.

Lalam: This work is important because it moves the needle away from using explanation quality as a proxy; it forces us to be explicit about what we are measuring when we claim a model is fair or trustworthy in legal contexts.

Tom: Fantastic summary, everyone. It sounds like we have a really clear direction now on how to approach fairness and explainability in these complex models. We’ve got some serious material here for the listeners who want to understand the real implications of this research on deploying AI responsibly in law and beyond.

More episodes

← Home