Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
cs.CY, cs.AI, cs.CL
Submitted: 2026-04-21
Updated: 2026-09-07
Comments: Second version, 45 pages, EMNLP 2026
Code: https://github.com/speedyapply/JobSpy
License: http://creativecommons.org/licenses/by/4.0/
The gist: Research has documented LLMs' name-based bias in hiring and salary recommendations.
Terminology
Abstract
Research has documented LLMs' name-based bias in hiring and salary recommendations. In this paper, we instead consider a setting where LLMs generate candidate summaries for downstream assessment. In a large-scale controlled study, we analyze nearly one million resume summaries produced by 4 models under systematic race-gender name perturbations, using synthetic resumes and real-world job postings. By decomposing each summary into resume-grounded factual content and evaluative framing, we find that factual content remains largely stable, while evaluative language exhibits subtle name-conditioned variation concentrated in the extremes of the distribution, especially in open-source models. Our hiring simulation demonstrates how evaluative summary transforms directional harm into symmetric instability that might evade conventional fairness audit, highlighting a potential pathway for LLM-to-LLM automation bias.
Sources
- GPT-4 Technical Report
- Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
- Prompt Engineering Techniques for Mitigating Cultural Bias Against Arabs and Muslims in Large Language Models: A Systematic Review
- A Systematic Study of Bias Amplification
- Prompting Fairness: Integrating Causality to Debias Large Language Models
- Bias Amplification in Artificial Intelligence Systems
- A Closer Look at System Prompt Robustness
- The Social Cost of Intelligence: Emergence, Propagation, and Amplification of Stereotypical Bias in Multi-Agent Systems
- Gemma 2: Improving Open Language Models at a Practical Size
- Qwen2.5 Technical Report
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework