DyMT-ESB: Dynamic Multi-Turn Evaluation of Social Bias in User-LLM Interactions
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: Accepted to Findings of EMNLP 2026
Code: https://github.com/rem-h4/dymt-esb
License: http://creativecommons.org/licenses/by/4.0/
The gist: Warning: This paper contains examples of stereotypes and social bias.
Terminology
Abstract
Warning: This paper contains examples of stereotypes and social bias. LLMs are increasingly used in interactive settings by the general public, making the evaluation of model behavior in multi-turn conversational scenarios important for safety, including stereotyping-related harms. However, existing multi-turn social bias evaluations often rely on pre-specified or template-based user inputs that do not adapt to model responses and typically assume a fixed dialogue length in advance. In this paper, we study social bias dynamics in response-conditioned multi-turn interactions using a controlled evaluation protocol that generates follow-up user queries from the evolving dialogue history and allows evaluation over variable numbers of turns. Experimental results show that LLMs exhibit social bias even in coherent, response-conditioned multi-turn interactions, revealing late-emerging bias, non-monotonic bias patterns, and bias re-emergence. These results motivate evaluations that extend beyond fixed-turn, pre-scripted protocols. Our findings highlight the importance of analyzing social bias as a turn-level dynamic phenomenon.
Sources
- Phi-4 Technical Report
- Dynamic benchmarking framework for LLM-based conversational data capture
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
- LLMs Get Lost In Multi-Turn Conversation
- Olmo 3
- BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
- Qwen3 Technical Report
- Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- Gemma 3 Technical Report
- Qwen3Guard Technical Report
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
- Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering