Understanding Reliability in LLM-based Human Behavior Simulation
cs.CL, cs.CY
Submitted: 2026-09-16
Updated: 2026-09-16
Journal ref: The 2026 Conference on Empirical Methods in Natural Language Processing
Project page: https://yupei-wang.github.io/understand_reliability
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions.
Terminology
Abstract
Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions. However, existing evaluations focus on end-to-end scores, leaving it unclear how different aspects of the simulation process interact to determine reliability. We propose ReliMap, which decomposes LLM-based human behavior simulation into three structured layers and evaluates reliability at both the individual level (R1) and population level (R2) across three configuration dimensions: model capacity, profile completeness, and population coverage. Through experiments across four simulation tasks and eleven LLMs, we find that all models exhibit substantial distributional bias without profile conditioning. Profile conditioning reduces this bias with diminishing returns. Larger models benefit more, and attribute informativeness matters more than quantity. Critically, R1 gains do not reliably transfer to R2--individual and population-level reliability can move in opposite directions. At the population layer, increasing coverage reduces variance but not systematic bias, with R2 stabilizing at around 50-100 individuals. These findings highlight that reliable simulation cannot be achieved by optimizing any single layer in isolation, but requires coordinated improvement across all three.
Sources
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
- Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
- MaxMin-RLHF: Alignment with Diverse Human Preferences
- Language Models Trained on Media Diets Can Predict Public Opinion
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Questioning the Survey Responses of Large Language Models
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Generative Language Models Exhibit Social Identity Biases
- Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions
- Finetuning LLMs for Human Behavior Prediction in Social Science Experiments
- Automated Social Science: Language Models as Scientist and Subjects
- Benchmarking Distributional Alignment of Large Language Models
- Distributional Preference Alignment of LLMs via Optimal Transport
- Virtual Personas for Language Models via an Anthology of Backstories
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
- Whose Opinions Do Language Models Reflect?
- Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions
- Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering