WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities
cs.CL
Submitted: 2026-09-02
Updated: 2026-10-06
Comments: under review, dataset available via https://github.com/jerryspan/WinoQueer-NL/
Code: https://github.com/jerryspan/WinoQueer-NL
License: http://creativecommons.org/licenses/by/4.0/
The gist: While English language models have been widely examined for anti-queer bias, Dutch models remain understudied.
Terminology
Abstract
While English language models have been widely examined for anti-queer bias, Dutch models remain understudied. To address this gap, we developed a culturally and linguistically adapted Dutch dataset based on the English WinoQueer benchmark, containing pairs of stereotypical and counter-stereotypical sentences. To validate and expand it, we conducted an online survey with 43 Dutch queer participants, confirming 145 of 171 stereotypes as culturally relevant and identifying 22 new biases through free-text responses. The final released dataset, comprising 42,906 sentences, was evaluated using a range of Dutch-specific and multilingual models, including both masked language models (MLMs) and autoregressive language models (ARLMs), with bias measured via a score comparing log-likelihoods of stereotypical versus counter-stereotypical sentences. While the mean bias score across models appeared neutral (50%), closer analysis revealed significant disparities: some models favored stereotypical sentences up to 97% of the time for transgender identities, but only 6% of the time for gay-related pairs, with transgender and non-binary identities consistently receiving the highest bias scores. Our findings highlight the importance of culturally grounded datasets for evaluating and mitigating biases that disproportionately impact marginalized groups in Dutch language models.
Sources
- BERTje: A Dutch BERT Model
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Queer People are People First: Deconstructing Sexual Identity Stereotypes in Large Language Models
- MBBQ: A Dataset for Cross-Lingual Comparison of Stereotypes in Generative LLMs
- Fietje: An open, efficient LLM for Dutch
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering