Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs
cs.CL, cs.AI
Submitted: 2024-07-15
Updated: 2026-05-09
Comments: v6: Updated title; LangFair repository: https://github.com/cvs-health/langfair
Journal ref: Proceedings of the Sixth Workshop on Language Technology for Equality, Diversity, Inclusion (LT-EDI 2026), Association for Computational Linguistics, pp. 10-26, 2026
Code: https://github.com/cvs-health/langfair
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Bias and fairness risks in Large Language Models (LLMs) vary substantially across deployment contexts, yet existing approaches lack systematic guidance for selecting appropriate evaluation metrics.
Terminology
Abstract
Bias and fairness risks in Large Language Models (LLMs) vary substantially across deployment contexts, yet existing approaches lack systematic guidance for selecting appropriate evaluation metrics. We present a decision framework that maps LLM use cases, characterized by a model and population of prompts, to relevant bias and fairness metrics based on task type, whether prompts contain protected attribute mentions, and stakeholder priorities. Our framework addresses toxicity, stereotyping, counterfactual unfairness, and allocational harms, and introduces novel metrics based on stereotype classifiers and counterfactual adaptations of text similarity measures. We release an open-source Python library, langfair, for practical adoption. Extensive experiments on use cases across five LLMs and five prompt populations demonstrate that fairness risks cannot be reliably assessed from benchmark performance alone: results on one prompt dataset likely overstate or understate risks for another, underscoring that fairness evaluation must be grounded in the specific deployment context.
Sources
- The Impossibility of Fair LLMs
- AI Fairness 360: An Extensible Toolkit for Detecting, Understanding, and Mitigating Unwanted Algorithmic Bias
- Fairness in Recommendation Ranking through Pairwise Comparisons
- Identifying and Reducing Gender Bias in Word-Level Language Models
- Fairness Through Awareness
- WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
- Bias and Fairness in Large Language Models: A Survey
- Intrinsic Bias Metrics Do Not Correlate with Application Bias
- Equality of Opportunity in Supervised Learning
- Reducing Sentiment Bias in Language Models via Counterfactual Evaluation
- Wasserstein Fair Classification
- Panda LLM: Training Data and Evaluation for Open-Sourced Chinese Instruction-Following Large Language Models
- Collecting a Large-Scale Gender Bias Dataset for Coreference Resolution and Machine Translation
- Fairness in Recommendation: Foundations, Methods and Applications
- Holistic Evaluation of Language Models
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- BBQ: A Hand-Built Bias Benchmark for Question Answering
- A Comprehensive View of the Biases of Toxicity and Sentiment Analysis Methods Towards Utterances with African American English Expressions
- Learning Optimal Fair Scoring Systems for Multi-Class Classification
- Aequitas: A Bias and Fairness Audit Toolkit
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering