Mitigating Social Desirability Bias in Random Silicon Sampling

arXiv:2512.22725 · cs.CL, cs.CY · Submitted 2025-12-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Mitigating Social Desirability Bias in Random Silicon Sampling".

Jane: The gist: Question reformulation most effectively improves alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We're moving on now to the title and authors of "Mitigating Social Desirability Bias in Random Silicon Sampling," and what that actually means for how we use these language models today.

Jane: The paper is testing if we can make AI-generated survey responses look more like actual human opinions by fixing that social desirability bias, which is when the AI just gives us answers that sound socially correct instead of reflecting a real mix of people.

Lu: They used data from the American National Election Study two thousand twenty to build these synthetic populations, where each synthetic person has a demographic profile matching a real one.

Meng: So they're essentially creating digital stand-ins for people across eight different characteristics—age, location stuff—just to see how those profiles interact with an AI when it's asked questions.

Tom: And then they test different ways of asking the questions to see which ones actually make the AI behave like a real person in that simulation.

Jane: The main finding here is that simple question reformulation works really well for reducing that bias, especially when dealing with sensitive or political topics, because it pulls the answers away from just being socially approved.

Lu: They tested five different ways of rephrasing things, including directly replicating old methods and trying stuff like reverse-coding or adding a preamble encouraging sincerity to see what works.

Tom: And they showed that one specific strategy, the reformulated condition, was consistently better across all seven demographic groups they studied for social and political questions in their silicon sampling study.

Jane: That means if you’re using AI to simulate public opinion or market research, changing the way you phrase your survey questions is a powerful way to get data that actually looks like what real people would say.

Meng: It’s practical because it means less time spent cleaning up skewed data later on; the simulation is more accurate from the start when you're building these digital stand-ins.

Tom: But they also pointed out some caveats, which is always important in research, saying this reformulation doesn't fix everything and some sensitive topics still cause concentration on those safe answers for smaller models.

Jane: And they also showed that other tricks, like priming or adding a preamble to tell the AI to be "truthful," didn't help much and sometimes actually made the answers more uniform.

Lu: They also looked at how the AI itself generates text—the decoding temperature—and found that slightly increasing that randomness helped improve things in their baseline setup.

Tom: So, for those of you listening who just want to know what this means for your day-to-day, it's a reminder that when we use these powerful language models to model people, we have to actively engineer the way we ask the questions.

Jane: And that leads us right into what they didn't cover—the limitations of this work. They mentioned that their study is limited by which specific demographic variables they looked at and that if an AI defaults to a safe answer, it could be because of its general training, not just a lack of knowledge about a specific group.

Meng: So the paper confirms that designing better prompts is a big part of the solution for getting useful silicon samples out there for science.

Tom: Right, and this whole idea—using AI to simulate populations—opens up some huge questions about how we trust these digital stand-ins when they're making decisions.

Jane: Next up, we’re going to look at what other papers are doing in different areas of AI research that tackle similar issues.

The paper's summary: Tom: We just talked about how question reformulation is the best way to fix social desirability bias in silicon samples, and now we’re looking at what else they suggested for improving those responses beyond just that one trick.

Jane: They weren't just sticking to one trick; they looked at a few different ways to tweak the prompt, and they found that some of these other methods actually gave them some bumps on the road too.

Lu: They explored reverse-coding, which means writing questions in a way that doesn't change what’s being asked but changes how it sounds, and they found those were really helpful for specific things like Race Diversity questions.

Meng: So it wasn't just one thing that worked; there were targeted approaches too depending on the topic you’re looking at, which is interesting for practical engineering.

Tom: Right, and then they looked at priming and preamble prompts, which are basically instructions to get the AI to be more thoughtful or sincere before it even answers.

Jane: And they found that those explicit instructions actually had mixed results; sometimes they made the AI give more uniform answers, which wasn't what they were aiming for when wanting diversity.

Lu: The main point they’re pushing is that neutral phrasing, that third-person approach, really lowers the pressure on the model to just pick a socially "correct" answer.

Tom: So if you’re building an AI system for social science research or anything that needs to simulate people, the paper suggests focusing your prompt engineering energy on making the wording as neutral and objective as possible.

Jane: It changes how we think about training these systems; it’s not just about feeding them data, it's about designing the interaction with them so they reveal more of their actual reasoning process.

Meng: From an engineering standpoint, that means our prompt templates need to be highly flexible and adaptable based on the kind of data we're trying to gather for our simulations.

Tom: And they highlighted a limitation again—it doesn't eliminate bias entirely; if the topic is deeply ingrained in culture or politics, those concentrated responses will still happen.

Jane: So while reformulation is strong, we still have to be aware that some systemic biases in the data or the model itself can create roadblocks no matter how good our prompt wording is.

Lu: It’s a call for ongoing refinement; as we build better LLMs, we need better ways to guide them past those built-in social filters.

Tom: Exactly. This whole study shows us that the work of mitigating bias isn't finished; it’s an ongoing process of careful design and testing when using silicon sampling.

Jane: And if you want to see how this applies to more complex situations, we’ll be looking at papers that deal with massive amounts of video data next.

The paper's improvements: Tom: So to wrap up on "Mitigating Social Desirability Bias in Random Silicon Sampling," the main improvement they show is that question reformulation is definitely the most reliable way we have right now to get silicon samples closer to real human data when we’re dealing with surveys.

Jane: That's the main takeaway—neutral phrasing reduces that pressure on the AI and gives us distributions that look more like what people actually think across demographics in their silicon samples.

Lu: It really suggests that prompt design is a huge lever; even small changes in wording can significantly shift how the AI responds to tricky questions when it’s dealing with things like Race Diversity.

Meng: For practical work, this means we should probably prioritize testing those reformulation techniques when we’re trying to simulate opinions across different groups because it makes the simulation much more useful for policy analysis.

Lalam: I think this is important because if we can get silicon samples that actually reflect diverse opinions, it helps build a more nuanced and less biased cultural understanding in the future.

Tom: Exactly, Lalam. It moves us closer to having digital personas that aren't just repeating the loudest voices in the data.

Jane: And while they didn't solve everything, they gave us a clear direction on how to tackle one of the biggest hurdles in using LLMs for social research without getting totally stuck in safe answers.

Lu: It’s a good piece of work because it shows that we can systematically explore these prompt engineering techniques to make the AI behave more like a representative human population in simulations.

Meng: I think the limitation they pointed out about model dependence is key—we have to remember that every LLM does this differently, which means we can't just assume one solution works for all of them.

Tom: True. So, for anyone interested in how we get these AI simulations to be more accurate and less biased, checking out this paper on "Mitigating Social Desirability Bias in Random Silicon Sampling" is definitely worth your time.

Jane: And next time, we’re going to switch gears completely and look at how video data is exploding right now with papers like those from THYME.

Conclusion: Tom: So we're wrapping up on "Mitigating Social Desirability Bias in Random Silicon Sampling," which basically shows that question reformulation is the best tool we have right now for getting silicon samples that actually look like real human responses, especially on sensitive topics.

Jane: That’s the main point—neutral phrasing reduces that pressure on the AI and gives us distributions that look more like what people actually think across demographics in their silicon samples.

Lu: It really suggests prompt design is a huge lever; even small changes in wording can significantly shift how the AI responds to tricky questions when it’s dealing with things like Race Diversity.

Meng: For practical work, this means we should probably prioritize testing those reformulation techniques when we’re trying to simulate opinions across different groups because it makes the simulation much more useful for policy analysis.

Lalam: I think this is important because if we can get silicon samples that actually reflect diverse opinions, it helps build a more nuanced and less biased cultural understanding in the future.

Tom: Exactly, Lalam. It moves us closer to having digital personas that aren't just repeating the loudest voices in the data.

Jane: And while they didn't solve everything, they gave us a clear direction on how to tackle one of the biggest hurdles in using LLMs for social research without getting totally stuck in safe answers.

Lu: It’s a good piece of work because it shows that we can systematically explore these prompt engineering techniques to make the AI behave more like a representative human population in simulations.

Meng: I think the limitation they pointed out about model dependence is key—we have to remember that every LLM does this differently, which means we can't just assume one solution works for all of them.

Tom: True. So, for anyone interested in how we get these AI simulations to be more accurate and less biased, checking out this paper on "Mitigating Social Desirability Bias in Random Silicon Sampling" is definitely worth your time.

Jane: And next time, we’re going to switch gears completely and look at how video data is exploding right now with papers like those from THYME.

Sashank Chapala, Maksym Mironov, Songgaojun Deng

Eindhoven University of Technology

cs.CL, cs.CY

Submitted: 2025-12-27

Updated: 2026-10-03

Importance score: 79/100

The gist: The gist: Question reformulation most effectively improves alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES.

Key concepts

Silicon Sampling
This involves using LLMs, conditioned on specific demographic profiles, to simulate entire human populations for research purposes. It aims to overcome limitations like high costs and small sample sizes associated with traditional human surveys.
Social Desirability Bias (SDB)
This is the tendency of LLMs to generate answers that are socially approved or 'safe' rather than being demographically representative. It occurs because models learn from data that favors certain social norms, leading them to avoid controversial or non-standard responses.
Question Reformulation
This strategy involves rewriting survey questions using neutral, third-person phrasing to reduce the perceived pressure on the LLM to give a specific opinion. This technique was found to be the most successful at minimizing SDB and achieving more diverse, representative answers.

Terminology

Summary

The gist: Question reformulation most effectively improves alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES.

Introduction

Large Language Models (LLMs) can simulate human emotions and opinions, from subjective labeling of Twitter posts (Törnberg, 2023; Yang et al., 2024) to producing behavior changes consistent with personality frameworks (Serapio-García et al., 2023; Besta et al., 2025). This led researchers to explore whether LLMs can be used to simulate entire populations for social sciences (Argyle et al., 2023; Yang et al., 2024); or marketing research studies Sarstedt et al. (2024). Using LLM-simulated respondents in these setups could help tackle the limitations of large sample sizes, high costs, and long execution times associated with human respondents (Sun et al., 2024). This led to the idea of silicon sampling, which refers to the use of LLM-generated agents with demographic conditioning to simulate population-level survey responses (Argyle et al., 2023; Sun et al., 2024).

Prior Work and Research Gap

Prior work found remarkable alignment between silicon and human samples on some topics; however they diverged more sharply when sensitive topics or groups were involved Sun et al. (2024). This divergence likely reflects Social Desirability Bias (SDB), i.e., the tendency of LLMs to generate socially approved rather than demographically representative answers (Salecha et al., 2024) (see Section 2). Although models have been successfully trained to avoid displaying any explicit stereotypes, implicit biases can remain embedded in internal representations and emerge only under indirect or carefully crafted queries Bai et al. (2025); Zhao et al. (2025). This suggests that careful prompt design White (2023) may reduce the social desirability bias and achieve more representative responses. The main goal of this study lies in the systematic exploration of prompt engineering techniques for reducing the social desirability bias in large population silicon sampling Sun et al. (2024).

Methodology

The methodology follows a four–stage pipeline: (1) extracting demographic distributions from the ANES 2020 dataset; (2) generating a synthetic (silicon) population by conditioning LLMs on these demographic profiles; (3) collecting survey responses under five prompt conditions; and (4) evaluating alignment between silicon and human responses using divergence-based metrics. To generate synthetic respondents, the researchers estimate empirical marginal distributions over K = 8 demographic variables from the ANES dataset, and then construct a silicon sample where each synthetic respondent is defined by a demographic profile sampled independently from these empirical marginals Ri = 1 with N = 5,441 (Sun et al., 2024).

The study employed three LLMs: the open-source Llama3.1-8B-Instruct (Llama-8B), Llama-3.1-70B-Instruct (Llama-70B), and the closed-source GPT4.1-mini Achiam et al. (2023). The four prompt design strategies tested were:

  1. Replicate Condition (0): Replicating prior work Sun et al. (2024) with unchanged prompts as a baseline;

  2. Reformulated Condition (1): Attempting to minimize SDB by reducing the perception of being evaluated or asked for an explicit opinion through neutral, third-person phrasing;

  3. Reverse-coded Condition (2): Including reverse-coded versions where this does not substantively alter the semantic meaning of the item;

  4. Priming Condition (3): Using conditioning prompts to create a more “Thinking” agent, prioritizing reasoning over emotion; and

  5. Preamble Condition (4): Adding a preamble encouraging sincere answers and promising no judgment.

Experimental Results

The results demonstrate that reformulated prompts most effectively improve alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES (Chapala et al., 2025). Specifically, for Llama-8B, reformulation improved alignment on 9 out of 10 questions and consistently yielded greater response diversity (Chapala et al., 2025). The Reformulated condition consistently reduces JS-divergence for all seven demographic groups on the social and political topics, Race Diversity and Refugee Allowing (Chapala et al., 2025).

In contrast, Reverse-coded Condition produced mixed results across eligible items; it achieved the greatest overall improvements in specific questions like Race Diversity (e.g., −0.151 to −0.193) and meaningfully improved alignment on Refugee Allowing and Income Inequality for most groups (Chapala et al., 2025). Priming and Preamble encouraged response uniformity and showed no systematic benefit for bias mitigation, often increasing JS-divergence, suggesting that explicitly instructing models to be “truthful” or “sincere” may inadvertently activate a perception of evaluation in the model (Chapala et al., 2025; Lynch et al., 2025).

Decoding Stochasticity and Robustness

The effect of decoding temperature was examined, showing that increasing the temperature from 0 to 1 reduces JS-divergence in the Replicate condition, indicating that a portion of SDB arises from low-entropy decoding (Chapala et al., 2025). Reformulated condition consistently outperforms all other prompting strategies at both temperatures, achieving the lowest JS-divergence at T=1 (0.0678) (Chapala et al., 2025). Furthermore, reformulation performs most consistently for policy and public safety questions (3/4), moderately for identity and social norms (2/3), and poorly for economic topics (1/3) across topic categories (Chapala et al., 2025). The effectiveness of question reformulation remains robust across temporal shifts in survey populations, suggesting it is a general mechanism for mitigating social desirability bias in LLM-generated survey responses, not tied to a specific survey snapshot (Chapala et al., 2025).

Conclusion

Our results identify question reformulation as the most effective and consistent mitigation strategy. Neutral, third-person rephrasing reduces evaluative pressure in question wording, consistently leading to more diverse response distributions (Chapala et al., 2025). Reformulation does not fully eliminate SDB; politically sensitive or culturally entrenched topics continue to elicit concentrated responses, particularly for smaller models (Chapala et al., 2025). Other strategies show limited or inconsistent benefits, and priming and preamble often increase response uniformity (Chapala et al., 2025). The study underscores the importance of careful prompt design for LLM-based population simulations (Chapala et al., 2025).

Limitations

The limitations noted include model dependence, where variation in baseline silicon-sampling performance across LLMs reflects model-specific training data and architectures (Chapala et al., 2025). Population coverage is limited to selected demographic variables and U.S. population represented in the ANES data (Chapala et al., 2025). Finally, disentangling social desirability bias from insufficient population knowledge remains challenging, as a model’s tendency to default to a “safe” response may reflect normative pressure or lack of group-specific knowledge (Chapala et al., 2025).

Ethical Considerations

Silicon sampling raises ethical complications that deserve careful attention, requiring researchers to clearly disclose the synthetic nature of these data and label synthetic outputs wherever they appear (Chapala et al., 2025). Transparency in prompt design, model parameters, and demographic conditioning is essential when simulating perspectives from marginalized groups (Chapala et al., 2025). Participant privacy must also be protected by relying on aggregate data and avoiding the reconstruction of individual responses (Chapala et al., 2025).

References

Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (Achiam et al., 2023).

Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. Using large language models to simulate multiple humans and replicate human subject studies (Aher et al., 2023).

Meta AI. Meta-llama-3.1-8b-instruct (Meta AI, 2024).

Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples (Argyle et al., 2023).

Neeraj Arora, Ishita Chakraborty, and Yohei Nishimura. Ai–human hybrids for marketing research: Leveraging large language models (llms) as collaborators (Arora et al., 2025).

Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L Griffiths. Explicitly unbiased large language models still form biased associations (Bai et al., 2025).

Improvements for AI systems

  1. A core improvement is implementing question reformulation as a primary prompt engineering strategy for sensitive topics, as question reformulation most effectively improve alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES. This allows the AI system to generate responses that are statistically more representative of human populations by neutralizing evaluative language and adopting a third-person formulation.

  2. The system can be refined to use reverse-coded phrasing for specific questions where feasible, as this strategy showed significant gains, specifically stating that "Reverse-coded condition achieves the greatest overall improvements in these specific questions, substantially lowering divergence for all groups on Race Diversity (e.g., −0.151 to −0.193) and meaningfully improving alignment on Refugee Allowing and Income Inequality."

  3. The system should avoid using explicit evaluative or judgmental language when conditioning responses, as this was shown by Reformulated Condition (1) which applies several modifications to the phrasing of the survey questions while ensuring that their meaning remains unchanged to minimize SDB. This directly addresses the finding that reformulation improves alignment with human response distributions in random silicon sampling.

  4. The system should incorporate a decoding strategy based on higher stochasticity, as Higher-temperature decoding allows the model to sample from a broader set of plausible responses, mitigating model collapse toward socially acceptable options, and this strategy consistently outperforms all other prompting strategies at both temperatures for Reformulated condition.

  5. The system must be explicitly programmed to use a Preamble as a baseline instruction, even if it shows limited benefit, because the analysis suggests that Priming and Preamble often increase response uniformity, which helps in understanding the structural biases of LLMs under explicit instruction.

Sources

Related papers