Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
cs.CY, cs.AI, cs.CL, cs.LG
Submitted: 2024-07-29
Updated: 2026-08-30
Comments: Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society; code available at https://github.com/kyrawilson/Resume-Screening-Bias Revised 8/29/2026 to include errata description Revised 8/29/2026 to include description of errata
Code: https://github.com/kyrawilson/Resume-Screening-Bias
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Artificial intelligence (AI) hiring tools have revolutionized resume screening, and large language models (LLMs) have the potential to do the same.
Terminology
Abstract
Artificial intelligence (AI) hiring tools have revolutionized resume screening, and large language models (LLMs) have the potential to do the same. However, given the biases which are embedded within LLMs, it is unclear whether they can be used in this scenario without disadvantaging groups based on their protected attributes. In this work, we investigate the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection. Using that framework, we then perform a resume audit study to determine whether a selection of Massive Text Embedding (MTE) models are biased in resume screening scenarios. We simulate this for nine occupations, using a collection of over 500 publicly available resumes and 500 job descriptions. We find that the MTEs are biased, significantly favoring White-associated names in 85.1% of cases and female-associated names in only 11.1% of cases, with a minority of cases showing no statistically significant differences. Further analyses show that Black males are disadvantaged in up to 100% of cases, replicating real-world patterns of bias in employment settings, and validate three hypotheses of intersectionality. We also find an impact of document length as well as the corpus frequency of names in the selection of resumes. These findings have implications for widely used AI tools that are automating employment, fairness, and tech policy.
Sources
- Identifying and Improving Disability Bias in GPT-Based Resume Screening
- Mistral 7B
- Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
- Generative Representational Instruction Tuning
- MTEB: Massive Text Embedding Benchmark
- Degendering Resumes for Fair Algorithmic Resume Screening
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- Debiasing Gender Bias in Information Retrieval Models
- Improving Text Embeddings with Large Language Models
- Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework