Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias
cs.CL, cs.CY
Submitted: 2026-01-26
Updated: 2026-08-29
Comments: Accepted to EMNLP 2026 (Main Conference)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Demographic cue-based evaluation is widely used to study how large language models (LLMs) adapt their responses to signaled demographic attributes within and across groups.
Terminology
Abstract
Demographic cue-based evaluation is widely used to study how large language models (LLMs) adapt their responses to signaled demographic attributes within and across groups. This approach typically relies on a single cue (e.g., names) as a proxy for group membership, implicitly treating different cues as interchangeable operationalizations of a single underlying identity-conditioned behavior. We test this assumption in realistic advice-seeking interactions spanning 14.8 million prompts, focusing on race and gender in a U.S. context. We find that cues for the same group induce only partially overlapping changes in model responses, yielding inconsistent conclusions about personalization, while bias conclusions are unstable, with both magnitude and direction of group differences varying across cues. We further show that these inconsistencies reflect differences in cue-group association strength and linguistic features bundled within cues that shape model responses. Together, our findings suggest that demographic conditioning in LLMs is not a cue-invariant category-level parameter but depends fundamentally on how identity is cued, reflecting responses to linguistic signals rather than stable demographic categories. We therefore call for multi-cue, mechanism-aware evaluations as a foundation for robust and interpretable claims about demographic variation in LLM responses.
Sources
- First-Person Fairness in Chatbots
- Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
- Language Models Change Facts Based on the Way You Talk
- Towards interactive evaluations for interaction harms in human-AI systems
- Evaluating and Mitigating Discrimination in Language Model Decisions
- IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
- The Need for a Socially-Grounded Persona Framework for User Simulation
- Race Discrimination in Internet Advertising: Evidence From a Field Experiment
- The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering