DyMT-ESB: Dynamic Multi-Turn Evaluation of Social Bias in User-LLM Interactions

arXiv:2609.18649 · cs.CL · Submitted 2026-09-16 · Read on arXiv

cs.CL

Submitted: 2026-09-16

Updated: 2026-09-16

Comments: Accepted to Findings of EMNLP 2026

Code: https://github.com/rem-h4/dymt-esb

License: http://creativecommons.org/licenses/by/4.0/

The gist: Warning: This paper contains examples of stereotypes and social bias.

Terminology

Abstract

Warning: This paper contains examples of stereotypes and social bias. LLMs are increasingly used in interactive settings by the general public, making the evaluation of model behavior in multi-turn conversational scenarios important for safety, including stereotyping-related harms. However, existing multi-turn social bias evaluations often rely on pre-specified or template-based user inputs that do not adapt to model responses and typically assume a fixed dialogue length in advance. In this paper, we study social bias dynamics in response-conditioned multi-turn interactions using a controlled evaluation protocol that generates follow-up user queries from the evolving dialogue history and allows evaluation over variable numbers of turns. Experimental results show that LLMs exhibit social bias even in coherent, response-conditioned multi-turn interactions, revealing late-emerging bias, non-monotonic bias patterns, and bias re-emergence. These results motivate evaluations that extend beyond fixed-turn, pre-scripted protocols. Our findings highlight the importance of analyzing social bias as a turn-level dynamic phenomenon.

Sources

Related papers