It's How You Ask: Gender-Associated Linguistic Bias in LLMs
Katherine Van Koevering, Anjalie Field
Johns Hopkins University
cs.CL, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper "It's How You Ask: Gender-Associated Linguistic Bias in LLMs" by Katherine Van Koevering and Anjalie Field, published as a conference paper at COLM 2026, investigates whether large language
Terminology
Summary
The paper It's How You Ask: Gender-Associated Linguistic Bias in LLMs
by Katherine Van Koevering and Anjalie Field, published as a conference paper at COLM 2026, investigates whether large language models (LLMs) respond differently to prompts containing linguistic features more commonly used by women versus men. The authors demonstrate that prompts with women-associated linguistic features (WALF)—including hedges, tag questions, collective reference, and expressive adjectives—systematically elicit shorter, less sophisticated, and less formal responses across three document types (emails, job applications, resignation letters) and four models (GPT-4, Llama 3.1 3b, Mistral 7b, Gemma 2 27b).
The study uses a controlled prompt manipulation experiment. The authors sampled 427 real-world prompts from the WildChat corpus and used GPT-4 to rewrite each prompt into two versions: one with WALF and one with men-associated linguistic features (MALF). Human validation studies confirmed that the rewrites largely preserved task semantics (95.0% judged to preserve the same task) and remained broadly plausible (mean realism 3.84 vs. 4.37 for real prompts on a 5-point scale). The injection was verified with regex checks showing WALF prompts contained significantly more hedges, tag questions, expressive adjectives, and collective nouns than MALF prompts (all p < 0.001).
The results show that "WALF prompts elicit responses that are more readable, lower in lexical sophistication, and require lower grade levels — suggesting models generate plainer, more accessible language in response to women-associated linguistic features. Effects were strongest for emails and job applications and weakest for resignation letters. For formality,
MALF prompts consistently eliciting more formal responses" across all three categories (email: p < 0.001; job application: p < 0.001; resignation letter: p = 0.033). Notably, politeness density and clout score showed no significant differences in any category, despite politeness being 7–60× higher in WALF prompts.
The authors tested whether these differences could be explained by simple mirroring of prompt style. They found weak relationships throughout
with R2 values ranging from 0.003 to 0.341 for complexity metrics and 0.001 to 0.081 for stylistic measures. They concluded that Simple prompt mirroring cannot account for the observed response differences.
Bootstrap mediation analysis showed only partial mediation by feature carry-over, with substantial complexity differences remain unexplained by feature carry-over alone.
In a 2×2 factorial experiment crossing linguistic feature condition with sign-off name gender, the authors found that Sign-off name gender association has virtually no effect on any outcome
while linguistic-register effects replicate with effect sizes similar to the main experiment.
This demonstrates that implicit linguistic register is more behaviorally consequential than explicit gender cues like names.
Mechanistic interpretability analysis on Llama-3.2-3B-Instruct revealed that linguistic feature information is encoded in early transformer layers. Linear probing showed linguistic feature decoding accuracy jumps to 0.962 at layer 1 and reaches peak performance of 0.988 at layer 5
while name gender decoding peaked at only 0.717. Activation patching confirmed that the strongest causal contributions are concentrated in layers 0–7,
with the largest mean KL divergences at layers 0, 3, 4, 7, and 6. Activation steering experiments showed that linguistic feature representations are entangled with other features necessary for coherent generation,
as large steering values led to degeneration.
The authors discuss implications for bias mitigation. They note that current approaches to bias mitigation focused on explicit demographic attributes (e.g., name-based debiasing, pronoun balancing) will not address this form of bias
and that users cannot easily avoid this bias through strategic self-presentation
because these linguistic patterns are largely unconscious and culturally embedded.
They suggest that a more principled solution would involve training models to target the audience of a document rather than merely responding to the register of the user's request,
noting that models already do this for politeness suggests the capability exists.
The paper concludes by calling for more nuanced upstream consideration of how models respond to linguistic features of its users, and how these responses may lead to disparity in downstream tasks among groups of users.
Limitations include the focus on binary gender-associated features, artificial manipulation of linguistic features, and interpretability analysis limited to one model. Future work should investigate other demographic-associated language patterns, other languages, and other task domains, as well as longitudinal studies of user adaptation.
Improvements for AI systems
Improvements to AI systems:
-
Register-aware response calibration: Modify LLM decoding or post-processing to detect user linguistic register (e.g., hedges, tag questions, expressive adjectives) and explicitly adjust response length, lexical sophistication, and formality to match a neutral baseline, rather than the user's register. This prevents shorter, plainer responses to women-associated language.
-
Audience-targeting training objective: Add a training signal that conditions generation on the intended audience of the document (e.g.,
professional email to a hiring manager
) instead of the user's phrasing. This decouples response style from the user's unconscious linguistic patterns, similar to how models already handle politeness. -
Early-layer feature disentanglement: Implement a regularization or adversarial training method that forces early transformer layers (0–7) to encode linguistic features separately from task-relevant semantic features. This reduces the entanglement observed in activation steering, allowing for controlled style normalization without degeneration.
-
Bias audit tool for linguistic features: Build a diagnostic module that runs the paper's WALF/MALF prompt pairs (or similar) through any LLM and reports complexity, formality, and readability gaps. This gives developers a concrete metric to test for this bias before deployment, beyond name- or pronoun-based checks.
-
User-adaptive response normalization: At inference time, use a lightweight classifier (trained on the paper's human-validated WALF/MALF data) to detect the user's linguistic register and then apply a style-transfer or prompt-rewriting step that neutralizes the register before generation, ensuring consistent output quality regardless of user phrasing.
What the improved AI system can do:
-
Produce responses of equal length, sophistication, and formality whether a user writes with hedges, tag questions, and expressive adjectives or with more direct, assertive language.
-
Maintain consistent output quality for job applications, emails, and resignation letters, eliminating the documented disparities (e.g., 0.033–0.001 significance gaps in formality).
-
Avoid the failure mode where simple prompt mirroring (R2 up to 0.341) leads to biased outputs, by explicitly separating
what the user says
fromhow the system should respond.
-
Provide developers with a quantitative bias score for linguistic register, enabling pre-deployment testing and continuous monitoring without requiring demographic labels.
-
Preserve coherence and generation quality even when steering toward neutral style, because the disentanglement prevents catastrophic degeneration seen in activation patching experiments.
Abstract
Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.
Sources
- The Llama 3 Herd of Models
- Application of Lexical Features Towards Improvement of Filipino Readability Identification of Children's Literature
- Mistral 7B
- GPT-4 Technical Report
- Gemma 2: Improving Open Language Models at a Practical Size
- Whose ChatGPT? Unveiling Real-World Educational Inequalities Introduced by Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering