A Heuristic Perspective on Debiasing Language Models
Tian Lan, Yemin Wang, Chuancheng Shi, Xiangyu Wu, Zesheng Shi, Yuan Wang, Jiang Li, Guanglai Gao, Xiangdong Su
cs.CL
Submitted: 2026-08-01
Comments: 13 pages in total, 5 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language models (LMs) often acquire various biases during pre-training and may express them in interactions, potentially causing social harm.
Terminology
Abstract
Language models (LMs) often acquire various biases during pre-training and may express them in interactions, potentially causing social harm. Existing methods often rely on counterfactual augmentation or representation projection. These strategies remain limited in practice due to their high computational costs and difficulty in scaling to larger models. Additionally, many of these strategies require manual data annotation, narrowing their scope to specific cultures and bias categories. To overcome these limitations, we propose HEIMAT, a HEurIstic-style autoMATic debiasing framework for LMs. HEIMAT consists of two main steps: bias disclosure and debiasing fine-tuning. In the first step, it uses simple templates to construct heuristic prompts, which are applied to reveal model biases and generate corresponding context prompts. In the second step, it fine-tunes the model by minimizing the Jensen-Shannon divergence of predictions on these context prompts to reduce bias. Extensive experiments show that HEIMAT effectively mitigates bias in different cultures while maintaining the model's natural language understanding (NLU) performance.
Sources
- Disclosure and Mitigation of Gender Bias in LLMs
- Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities
- DNA: Uncovering Universal Latent Forgery Knowledge
- C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
- Bias and Fairness in Large Language Models: A Survey
- Language (Technology) is Power: A Critical Survey of "Bias" in NLP
- On the Usability of Transformers-based models for a French Question-Answering task
- Fast Model Debias with Machine Unlearning
- Self-Debias: Self-correcting for Debiasing Large Language Models
- TinyBERT: Distilling BERT for Natural Language Understanding
- Debiasing Pre-trained Contextualised Embeddings
- Common to Whom? Regional Cultural Commonsense and LLM Bias in India
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- Towards Debiasing Sentence Representations
- DeepSeek-V3 Technical Report
- Decoupled Weight Decay Regularization
- Gender Bias in Neural Natural Language Processing
- On Measuring Social Biases in Sentence Encoders
- An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models
- Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering