When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "When Synthetic Users Fail".
Jane: Large language models (LLMs) are increasingly used as synthetic users to generate evidence for product, policy, and market decisions, but this substitution can be invalid if not rigorously evaluated.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to wrap up, the paper "When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses" really drives home that we can't just trust the outputs from these LLMs without checking for these specific flaws.
Jane: They give us a concrete benchmark and an evaluation framework that tells practitioners whether synthetic user evidence is trustworthy before they use it for a decision at hand <ref:2607.26348#pg1>.
Lu: The contribution is building that cross-domain benchmark itself, packaged as a reusable toolkit that anyone can use to test their own synthetic user systems <ref:2607.26348#pg0>.
Meng: From an engineering standpoint, the idea of a baseline-anchored evaluation, where you compare against non-LLM baselines to see if there’s any individual advantage at all, that's a practical tool we can actually build into our pipelines <ref:2607.26348#pg1>.
Lalam: And they introduced this stereotyping index, which is a single number meant to compare how predictive an attribute is in the model versus in real people <ref:2607.26348#pg2>.
Tom: The implication for us is that we need to check for individual fidelity separately from aggregate fidelity, and we have to be careful about the exaggeration of demographic influence <ref:2607.26348#pg1>.
Jane: It’s a warning that simply using a bigger model isn't the solution; you still need to look at how those models handle individual answers and stereotyping, which is what this work on synthetic users does <ref:2607.26348#pg2>.
Conclusion: Tom: So, we're wrapping up on this paper, "When Synthetic Users Fail." Basically, they’re showing us that when we use these large language models to act like real people for surveys or decisions, there are two major ways those simulations fall apart.
Jane: Two distinct failures that show up no matter what the model is or which domain you're in. It’s about whether the AI can actually capture what makes a human answer unique versus just repeating general trends.
Lu: The authors set up this comparison using real data from things like U.S. social attitudes and cross-cultural values, testing them across different models and different sizes of those models too.
Meng: And they found that the issue isn't just about accuracy in general; it’s more specific—it points out a problem with stereotyping where the AI tends to exaggerate how much one thing predicts an answer compared to what actually happens in real people.
Lalam: From my side, I see this as a critical test for how we use synthetic data for policy or market research because if the model is over-determined, you get decisions based on made-up human behavior patterns.
Tom: Exactly. So when we look at the title and who wrote it, this isn't just another paper about LLM performance; it's a framework designed to make sure that synthetic user evidence actually holds up under scrutiny.
Jane: It’s less about whether the AI can sound human and more about whether its simulated answers are actually reliable for making real-world decisions.
Lu: The contribution here is giving us a way to test these systems rigorously, not just by asking them questions, but by comparing their results against solid human baselines that aren't AI generated.
Meng: And the framework they propose is pretty practical—it tells people exactly what checks they need to run before they let these synthetic users guide any important choice.
Lalam: It really shifts the focus from "what can this model say?" to "is this simulation trustworthy for making a call?" and that’s a big deal for anyone building tools on top of generative AI.
Tom: So, we've seen the failures, we've seen the checks they suggest, and now we see how it all ties together—it’s about building better guardrails so these simulations don't mislead us.
Jane: It means that for decision support systems using synthetic users, you gotta look beyond just a high accuracy score and check those specific fidelity points.
Tom: And if you want to know exactly what those checks look like in practice, we’ve got some more deep dives into the validation framework coming up next.
Zihan Chen, Di Zhu, Lei Nico Zheng
Stevens Institute of Technology · University of Massachusetts Boston
cs.CL, cs.AI, cs.CY, cs.HC
Submitted: 2026-07-28
Updated: 2026-10-04
Comments: 19 pages, 5 figures, 4 tables. Preprint; under review. Code and data are publicly available at: https://github.com/ZihanChen1995/when-synthetic-users-fail-a-cross-domain-benchmark-of-llm-simulated-human-survey-responses
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: Large language models (LLMs) are increasingly used as synthetic users to generate evidence for product, policy, and market decisions, but this substitution can be invalid if not rigorously evaluated.
Key concepts
- Explicit Naive Demographic Baseline
- This is a benchmark created by fitting the conditional distribution of real human answers given specific demographics. It serves as a 'naive' predictor that captures the irreducible ceiling of what demographics alone can explain in survey responses, helping to judge if an LLM actually adds value beyond this baseline.
- Individual Fidelity vs. Trivial Baseline
- This concept tests whether an LLM simulating one person is actually more accurate than just using the demographic information alone. The paper found that under testing conditions, LLMs fail to provide this individual-level advantage over a simple demographic predictor.
- Demographic Over-determination (Stereotyping Index)
- This failure occurs when LLMs exaggerate the predictive power of demographics. For example, if politics explains only 1.5% of real answer variation, the model might falsely suggest it explains 67%. This means the model reinforces stereotypes rather than accurately reflecting real human behavior.
- Decision-Impact Analysis
- This is a practical check for using LLM simulations in business decisions. It quantifies how much models inflate differences between groups (like segment gaps) and how often they would send teams to the wrong group based on their biased predictions.
Terminology
Summary
Large language models (LLMs) are increasingly used as synthetic users to generate evidence for product, policy, and market decisions, but this substitution can be invalid if not rigorously evaluated. This paper develops a cross-domain benchmark and validation framework to determine when LLM-simulated human survey responses are trustworthy for decision support.
The gist: Under demographic prompting and the survey-simulation protocols we test, LLM synthetic users exhibit two distinct failures that replicate across both domains, all four models, and both families.
How it works
The study applies a single protocol across four models spanning two families and an 8B-to-frontier capability range to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey)<ref:2607.26348#pg2> Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families<ref:2607.26348#pg2>.
The evaluation is anchored by a explicit naive demographic baseline
: for each question we fit the conditional distribution of human answers given demographics on a held-out portion of the real data, and score every LLM against it on the same respondents<ref:2607.26348#pg2>. This yardstick is crucial because without it, “the model predicts individuals with X% accuracy” is uninterpretable, because individual survey answers are not a deterministic function of demographics; there is an irreducible ceiling that a trivial predictor already captures<ref:2607.26348#pg2>.
Research Questions
The study formalizes its investigation into five research questions:
-
RQ1 Individual fidelity vs. a trivial baseline: Do LLM synthetic users predict individual human answers more accurately than a naive demographic-conditional predictor?<ref:2607.26348#pg2>.
-
RQ2 Aggregate fidelity: Do LLM synthetic users reproduce the population-level distribution of human answers?<ref:2607.26348#pg2>.
-
RQ3 Subgroup structure: Do LLMs represent the demographic structure of attitudes faithfully, or do they distort how predictive demographics are of answers?<ref:2607.26348#pg2>.
-
RQ4 Stability and capability: Are these behaviors stable across model family, model capability, and output format?<ref:2607.26348#pg2>.
-
RQ5 Cross-domain transfer: Do the answers to RQ1–RQ4 hold in both domains, or are they domain-specific?<ref:2607.26348#pg2>.
Headline Findings
The central result is that under demographic prompting and the survey-simulation protocols we test, LLM synthetic users exhibit two distinct failures that replicate across both domains, all four models, and both families<ref:2607.26348#pg2>. The first failure is a lack of individual-level advantage: “In short, an LLM asked to role-play an individual adds no information beyond what the demographics alone already imply”<ref:2607.26348#pg2>. On GSS (U.S. attitudes) every LLM, at best, only ties the lookup table and trails the learned baseline; on WVS (cross-cultural values) every model is 11 to 22 percentage points less accurate than the baseline<ref:2607.26348#pg2>.
The second failure is demographic over-determination, or stereotyping: “Models consistently exaggerate it” when measuring how strongly a demographic attribute predicts a person’s answer<ref:2607.26348#pg2>. For U.S. political leaning and confidence in banks, for example, a person’s politics explains only about 1.5% of the variation in real answers, yet the model behaves as though it explains up to roughly 67%<ref:2607.26348#pg2>. Neither failure is fixed by using a bigger, more capable model: “the frontier models stereotype at least as strongly as the small 8B model, and often more”<ref:2607.26348#pg2>.
Contributions
The contributions are methodological and practical. First, we build a “compact, reproducible cross-domain benchmark for LLM synthetic users,” packaged as a reusable evaluation toolkit available on request<ref:2607.26348#pg2>. Second, we develop a “baseline-anchored evaluation,” showing that the standard individual-accuracy number is uninterpretable without non-LLM baselines, and that against them LLMs show no individual-level advantage<ref:2607.26348#pg2>. Third, we introduce a “stereotyping index,” a single, bounded number that compares how predictive a demographic attribute is of the answer in the model versus in real people<ref:2607.26348#pg2>. Fourth, we provide a “decision-impact analysis” that carries the over-determination finding through to the decision: on the canonical segment-targeting task, we quantify how far the models inflate between-segment gaps (two to fourfold), how often they would send a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people<ref:2607.26348#pg2>.
Validation Framework
The practical deliverable is a “validation framework for intelligent synthetic-user systems used in decision support.” This framework requires practitioners to run checks before deployment, including:
(1) reporting individual-level and aggregate-level fidelity separately<ref:2607.26348#pg2>.
(2) always benchmarking individual fidelity against non-LLM baselines computed on real held-out data using distance-aware or proper-scoring metrics on ordinal scales<ref:2607.26348#pg2>.
(3) reporting subgroup determinism (e.g., ∆η2), not just group means, to detect stereotyping<ref:2607.26348#pg2>.
(4) before acting on any segment-level read, checking the decision-impact quantities (gap inflation, wrong-target rate, and spurious-split rate) against held-out human data<ref:2607.26348#pg2>.
(5) reporting invalid/refusal rates, especially for distribution prompts and smaller models<ref:2607.26348#pg2>.
(6) not assuming a larger or more capable model is a safer synthetic user<ref:2607.26348#pg2>.
(7) restricting validity claims to the population, domain, and level of analysis actually tested<ref:2607.26348#pg2>.
The remainder of the paper reviews related work, details the data and protocol, presents results by research question, and discusses implications and limitations<ref:2607.26348#pg2>.
REFERENCES
[1] Marko Sarstedt, Susanne J Adler, Lea Rau, and Bernd Schmitt. Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines. Psychology & Marketing, 41(6):1254–1270, 2024.<ref:2607.26348#pg9>
[3] Galit Shmueli and Otto R Koppius. Predictive analytics in information systems research1. MIS quarterly, 35(3):553–572, 2011.<ref:2607.26348#pg10>
[4] Robert M O’Keefe and Daniel E O’Leary. Expert system verification and validation: a survey and tutorial. Artificial intelligence review, 7(1):3–42, 1993.<ref:2607.26348#pg10>
[5] Lisa P Argyle, Ethan C Busby, Nancy Fulda, Joshua R Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3):337–351, 2023.<ref:2607.26348#pg10>
[6] James Bisbee, Joshua D Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M Larson. Synthetic replacements for human survey data? the perils of large language models. Political Analysis, 32(4):401–416, 2024.<ref:2607.26348#pg10>
[7] Yuan Gao, Dokyun Lee, Gordon Burtch, and Sina Fazelpour. Take caution in using llms as human surrogates. Proceedings of the National Academy of Sciences, 122(24):e2501660122, 2025.<ref:2607.26348#pg10>
[8] Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. Using large language models to simulate multiple humans and replicate human subject studies. In International conference on machine learning, pages 337–371. PMLR, 2023.<ref:2607.
Improvements for AI systems
-
The improved system must incorporate a validation protocol to ensure trustworthiness before deployment, which includes: "(1) report individual-level and aggregate-level fidelity separately; (2) always benchmark individual fidelity against non-LLM baselines computed on real held-out data (at least a demographic lookup, ideally a learned demographic model) using distance-aware or proper-scoring metrics on ordinal scales."
-
The system should utilize the stereotyping index and its
coding-invariant companion measure
to detect whetherdemographics are more decisive of the answer than they are among real people,
treating this as a critical validity condition rather than a minor robustness check. -
For any decision-support task like segment targeting, the system must calculate and report decision-impact quantities:
(i) the between-segment gap maxg pg − ming pg; (ii) the wrong-target rate, the fraction of (question, axis) pairs where the model would send a team to the wrong segment; and (iii) how often it
manufactures a segment split that does not exist in real people." -
The system must adopt a strict rule:
do not assume a larger or more capable model is a safer synthetic user,
as capability does not fix the identified failures, including stereotype exaggeration. -
For outputs requiring probability estimates, the system should use
proper scoring rules
like log-loss and Brier score to ensure thatthe distribution puts mass where the human actually answered,
rather than relying solely on arg-max predictions which can be misleading.
Sources
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering