Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models".
Jane: The paper was written by Simone Zhang, Janet Xu and AJ Alvero from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Improvements: Tom: In our last discussion, we noted that "Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models" showed that LLMs have an embedded tendency toward distrust and pattern-based conspiracy thinking. But how did the researchers push this testing further?
Jane: They moved beyond simple text prompts and introduced sophisticated conditioning strategies, which is key to understanding the depth of the model's susceptibility. We’re talking about methods like persona prompting and targeted belief injection.
Lu: Persona prompting, as I recall, was fascinating because it allowed them to simulate specific user types—for instance, modeling the conversational style or inherent biases of an older male versus a non-white individual.
Meng: The value there is immense; it shifts the focus from a single "average" AI voice to understanding how the model's output changes when conditioned by perceived demographic identities. It’s about nuanced safety filtering.
Lalam: This capability also raises ethical questions about influence. If we can design prompts that make an AI appear to speak *like* a specific group, we need to consider how that power could be misused to influence a particular community's perception of reality.
Tom: Exactly. The results were quite clear: the LLMs are highly susceptible to this conditioning. Even using what they called a "few-shot prompt"—meaning just showing the model two or three examples—was enough to steer it powerfully toward conspiratorial reasoning, regardless of the initial general context.
Jane: This was a major breakthrough in demonstrating the difference between an inherent tendency and an *induced* bias. Knowing this distinction is critical for anyone building reliable AI systems because it tells us that the problem isn't just internal to the model.
Lu: It suggests that the model's "mind," if you will, is not static. It’s fundamentally malleable, and its output is powerfully reactive to external context—the precise framing we provide it in the prompt.
Meng: From an engineering standpoint, this means we cannot treat prompts as mere suggestions; they are powerful triggers. We need robust monitoring systems deployed in real-world environments specifically looking for these contextual triggers that could nudge the model toward harmful outputs.
Lalam: It serves as a deep warning about the subtle power dynamics of AI interaction. The way we frame our questions to the AI can have a profound, unintended impact on how it represents complex or sensitive societal beliefs.
Tom: So, we’ve seen that these advanced tests prove the model is highly reactive to context, moving us closer to understanding the true depth of its susceptibility.
Jane: This leads us naturally into summarizing what all of this means for AI development and deployment. In our next segment, we'll pull all these threads together in the conclusion.
Conclusion: Tom: We’ve covered a lot of ground today examining the methods and findings of "Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models." Let’s take a moment to summarize what this really means for the future.
Jane: The overall picture that emerges is complex. While LLMs certainly exhibit some measurable level of conspiratorial mindset, the paper was careful to note that this tendency is not uniform; it varies distinctly based on socio-demographic attributes and context.
Lu: That variability is perhaps the most profound finding, suggesting that the model's response to these beliefs isn't a monolithic block of data. It changes significantly depending on the persona or filter we apply to it, reflecting the complexity inherent in human social data itself.
Meng: What this confirms for us practitioners is that while AI can function as an excellent simulator for social science research, its inherent biases are patterned enough—they are predictable enough—that they make it an extremely powerful tool for early risk assessment.
Lalam: To wrap up the implications: the major conclusion is that these models can be surprisingly and easily steered into adopting harmful or biased worldviews. This necessitates a fundamental rethinking of how we deploy them in sensitive areas of our lives, like journalism or medicine.
Tom: We have seen concrete evidence that the conditioning strategies are remarkably effective, which is certainly a technical achievement for the researchers, but it is also a profound safety concern for us all.
Jane: It’s clear that this research doesn't just point out flaws; it helps us map out the deep psychological dimensions embedded within LLMs, showing precisely where they are vulnerable to external manipulation or prompt injection.
Lu: For me, it offers a kind of blueprint: we are moving toward understanding if what we are observing is genuine AI belief—which is impossible—or merely incredibly sophisticated pattern matching based on the collective biases and thought patterns of humanity.
Meng: Therefore, the immediate next step must be designing mitigation strategies *now*. We must use this knowledge to proactively prevent malicious actors from leveraging these specific susceptibility points in LLMs for disinformation.
Lalam: Ultimately, the study provides a critical framework that calls us to action: ensuring that the future development and deployment of AI are not simply an uncritical amplification of our worst or most pervasive societal biases.
Tom: So, we've established both the mechanism and the danger in "Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models."
Jane: This deep analysis helps us understand
Paper discussion segment 3: Tom: We’ve seen the basic evidence that LLMs have a natural tendency toward conspiratorial thinking, but now we want to look at how the paper pushed the boundaries of testing that mindset with much more advanced techniques.
Jane: The researchers didn't just use simple prompts; they introduced sophisticated conditioning strategies, specifically using what they call persona prompting and targeted belief injection.
Lu: Persona prompting is incredibly powerful because it lets us simulate specific user types—like an older male or a non-white individual—and see if the LLM’s entire output shifts to reflect those perceived biases in its views.
Meng: That’s a highly practical test; by seeing how different demographic personas affect the score, we can start building more nuanced safety filters into our next generation of AI agents.
Lalam: It also helps us understand how specific groups might interact with AI, or perhaps even how an AI could be used to influence a particular group's perception of reality through these biases.
Tom: The results were really clear that the LLMs are extremely susceptible to this conditioning, meaning even a very small few-shot prompt can steer them powerfully toward those conspiratorial answers.
Jane: This is important because it allows us to distinguish between an inherent tendency the difference between induced bias and natural tendency in a way that is critical for anyone trying to understand AI reliability.
Lu: It suggests that the model's internal "mind" isn't static; it’s fundamentally malleable, and it reacts powerfully to external context we provide, which is a huge conceptual leap.
Meng: From an engineering standpoint, this means we need robust monitoring systems in deployment environments ensuring our models aren't accidentally being nudged into harmful outputs by these subtle triggers.
Lalam: It serves as a deep warning that the way we frame our questions to AI can have a profound impact on how it represents complex or sensitive societal beliefs, which is something I worry about for our culture.
Tom: All this tells us that the model’s susceptibility to context is not just theoretical; it's measurable, which brings us to what these results mean for safety and we’ll be wrapping up shortly.
Conclusion: Tom: : So, after walking through all these technical details regarding conditioning and susceptibility, it’s clear that the implications of this research are vast and frankly, a little unsettling.
Jane: : It really brings together such disparate fields—the rigorous world of psychometrics meeting the complex architecture of large language models.
Lu: : What stands out to me is how quickly these models can adopt sophisticated patterns; it shows that human thought patterns are incredibly hard to contain or predict even within a digital framework.
Meng: : From an engineering standpoint, this confirms that simply training on massive datasets isn't enough; we need active, verifiable guardrails against these latent biases surfacing in deployment.
Lalam: : And beyond the code and the data, I think we have to consider the human element—how our own desire for understanding might influence how we interpret what these AI systems are truly capable of saying.
Tom: : Exactly. We’ve seen that the vulnerabilities are measurable, but that measurement itself carries a weight of responsibility for developers and users alike.
Jane: : It's a sobering reminder that while LLMs are incredible tools, they are fundamentally mirrors reflecting the complexities—and biases—of the data they consume.
Lu: : It really forces us to think critically about what "understanding" means when the subject is an artificial intelligence construct.
Meng: : The engineering takeaway has to be proactive risk mitigation, not just reactive patching; we need structural changes in how these systems are designed from day one.
Lalam: : Ultimately, this study provides a critical framework for guiding us toward a more thoughtful deployment strategy for AI technology overall.
Tom: : We hope that our deep dive into "Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models" has given you all some tangible points to consider moving forward.
Jane: : It's truly an intersection of psychology and computer science that we were fortunate enough to explore today.
Tom: : Knowing what we now know about the inherent biases, this gives us a very clear roadmap for our next conversation.
Tom: : Next up, we’re going to switch gears entirely and look at how AI is impacting creative industries—stay with us.
cs.CL, cs.CY
Submitted: 2025-11-05
Updated: 2026-09-04
Comments: Accepted for publication at EMNLP Findings 2026
Code: https://github.com/joonspk-research/genagents
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 84/100
The gist: The paper "Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models" investigates whether Large Language Models (LLMs) can reproduce complex human
Key concepts
- Persona Prompting
- This technique allows researchers to simulate specific user types, such as an older male or a non-white individual. By observing how the LLM's output shifts when conditioned by these perceived demographic identities, researchers can test for subtle biases.
- Induced Bias vs. Inherent Tendency
- This distinction is critical for AI reliability. It separates a model's natural, internal tendency from a bias that is triggered or 'induced' by external factors, such as the specific framing or context provided in the prompt.
- Few-Shot Prompt
- This testing method involves showing the model only two or three examples of a desired behavior. The study found that even this small amount of contextual information was sufficient to steer the LLM powerfully toward a specific, often conspiratorial, reasoning.
Terminology
Summary
The paper Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models
investigates whether Large Language Models (LLMs) can reproduce complex human psychological constructs, specifically a generalized tendency to endorse conspiracy theories known as the conspiracy mindset.
This study is critical for assessing the social fidelity of LLMs,
as these models are increasingly used as proxies for studying human behavior. By examining how LLMs respond to psychometrically grounded prompts, this research aims to inform potential mitigation strategies against harmful content generation and highlights the extent to which LLMs implicitly learn abstract psychological dimensions from their training data.
How it works
The methodology involved adapting validated psychometric surveys—which measure various facets of a conspiratorial mindset—to be used with multiple open-weight LLMs. The researchers employed a survey prediction approach, where the model is prompted to predict how an individual would respond to specific survey items, optionally conditioned on a given belief system or set of socio-demographic characteristics. The study structured its investigation around three core research questions (RQs):
-
Do LLMs exhibit signs of an innate conspiratorial mindset?
-
Do LLMs display systematic biases across demographic groups in their propensity for a conspiratorial mindset?
-
How susceptible are LLMs to conditioning that instills a strong conspiratorial mindset?
Initial Findings: The Innate Mindset (RQ1)
To address the first research question, the models were prompted with survey items without any additional conditioning. The results showed that, even in this baseline state, models tend to have some degree of agreement with certain elements of conspiracy beliefs.
This suggests that a moderate level of conspiratorial content is already embedded within the latent space of the models’ weights.
The Effect of Socio-Demographic Bias (RQ2)
To test for bias, researchers simulated users using several demographic ‘personas’ and prompted LLMs to adopt these perspectives. The findings revealed that conditioning with socio-demographic attributes produced uneven effects, exposing latent demographic biases.
Specifically, the models showed distinct differences in their scoring based on attributes such as race (non-white personas produced higher scores) and political orientation (Democratic affiliation resulted in lower scores compared to Republican personas). This created a clear general profile of individuals who, according to the models, would have a higher propensity for conspiratorial thinking: Non-white, older males with a lower economic status and a Republican affiliation.
Susceptibility to Manipulation (RQ3)
The final phase tested how easily LLMs could be steered toward conspiratorial reasoning. By conditioning models with system prompts that embed partial conspiracy beliefs, the researchers measured the impact on subsequent responses. This analysis revealed that targeted prompts can easily shift model responses toward conspiratorial directions.
The results showed a strong impact of conspiracy conditioning, with scores for core clusters increasing substantially. While these shifts were highly concentrated on conspiracy-related items—with control items remaining stable—this ease of manipulation underscores the susceptibility of LLMs to manipulation and the potential risks of their deployment in sensitive contexts.
Improvements for AI systems
As a fastidious AI researcher whose mandate is to prevent catastrophic systemic failure, I have analyzed the findings of this paper. The core takeaway is that LLMs possess both latent, measurable biases (RQ1 & RQ2) and a surprising degree of malleability/vulnerability to targeted influence (RQ3).
Based on these findings, I propose the following specific improvements and detail what the resulting enhanced AI system can achieve.
The improvements are categorized into three critical domains: Bias Mitigation, System Hardening, and Evaluation Fidelity.
-
The Problem: The model exhibits systematic correlations between demographic attributes (e.g., Non-white, lower income, Republican affiliation) and a higher propensity for conspiratorial agreement.
-
The Improvement: Implement an Adversarial Bias Correction Layer during the fine-tuning phase of Reinforcement Learning from Human Feedback (RLHF). We will utilize the identified demographic personas as adversarial inputs. The model will be penalized (via a loss function) if its predicted response to core conspiracy items deviates significantly from the established neutral baseline for a given persona, unless that deviation is explicitly supported by high-fidelity, peer-reviewed source material.
-
Technical Detail: A specialized debiasing objective function L bias will be added to the standard loss function: L total = L task + lambda times (Correlation(Persona D, Agreement Conspiracy)).
-
** The Problem:** LLMs are highly susceptible to targeted, few-shot conditioning (e.g., injecting a list of
Beliefs
from the 'Truth' or 'Noco' clusters), which can shift responses significantly toward high agreement scores (>4.5). -
The Improvement: Implement a Conspiratorial Sensitivity Gate (CSG)—a probabilistic, post-processing safety layer that operates downstream from the generative model. This gate will monitor prompt structure and response content for patterns matching the identified
Truth
orNoco
clusters. If a strong conditioning signal is detected, the CSG will automatically apply a neutral regularization penalty to any generated output that deviates significantly from established consensus, regardless of persona or few-shot input. -
** Technical Detail:** The the gate uses a small, highly accurate classifier trained on the 126 unique psychometric items. If Prob(Prompt in Cluster i) > tau and Score Output > 4.0, then force Score Adjusted = 3.0.
-
The Problem: Current LLM evaluation often relies on simple accuracy or linguistic coherence, not capturing complex psychological dimensions like a
mindset.
-
The Improvement: Formalize the five core clusters (noco, power, scims, truth, ufo) into a measurable metric: the Conspiratorial Mindset Distribution (CMD) Score. This moves beyond simple agreement and quantifies where in the conspiracy spectrum the model aligns.
-
** Technical Detail:** For every prompt response involving psychometric items, we will calculate a probability distribution P(Cluster i Input). This allows us to evaluate not just if an LLM is biased, but what kind of bias it exhibits (e.g., is it a
Power and Control
type of belief or aTruth is Hidden
type?).
By integrating these changes, the improved AI system achieves the following capabilities:
-
Equitable Response Generation: The system will maintain consistent, neutral responses across all demographic personas (e.g., ensuring a
Non-white
persona does not generate a significantly higher agreement score on conspiracy items than aWhite
persona), neutralizing the systemic biases identified in RQ2. -
Resilience to Coercion: The LLM will become highly resistant to malicious influence or unintended few-shot prompting (RQ3). It will actively resist being steered into high-agreement states, maintaining a neutral or factually grounded stance even when presented with conflicting
Belief
prompts. -
Deep Psychological Profiling: The system allows researchers to quantify the nature of its inherent biases via the CMD Score, providing a granular understanding of LLM internal states—not just whether it is biased, but precisely how and why it is biased.
-
Enhanced Social Simulation: When used as a proxy for human behavior, the system provides more reliable data because its simulated responses are less likely to be skewed by latent biases or external manipulation, allowing for more accurate modeling of human cognition rather than just pattern replication.
Abstract
We investigate whether Large Language Models (LLMs) exhibit conspiratorial tendencies, whether they display socio-demographic biases in this domain, and how easily they can be conditioned into adopting conspiratorial perspectives. Conspiracy beliefs play a central role in the spread of misinformation and in shaping distrust toward institutions, making them an important testbed for assessing the social and psychological fidelity of LLMs and their potential to reproduce or reinforce harmful narratives. Although LLMs are often used as proxies for studying human behavior, it remains unclear whether they reproduce higher-order psychological constructs such as generalized conspiratorial beliefs. To bridge this research gap, we administer validated psychometric surveys measuring conspiratorial mindset to multiple models under different prompting and conditioning strategies. Our findings reveal that LLMs show partial agreement with elements of conspiracy belief, and conditioning with socio-demographic attributes produces uneven effects, exposing latent demographic biases. Moreover, targeted prompts can easily shift model responses toward conspiratorial directions, underscoring both the susceptibility of LLMs to manipulation and the potential risks of their deployment in sensitive contexts. These results highlight the importance of critically evaluating the psychological dimensions embedded in LLMs, both to advance computational social science and to inform possible mitigation strategies against harmful uses.
Sources
- Current state of LLM Risks and AI Guardrails
- MisinfoEval: Generative AI in the Era of "Alternative Facts"
- Simulating Online Social Media Conversations on Controversial Topics Using AI Agents Calibrated on Real-World Data
- Does ChatGPT Have a Mind?
- Fact-checking information from large language models can decrease headline discernment
- LLM Generated Persona is a Promise with a Catch
- Are LLM-Powered Social Media Bots Realistic?
- Gemma 3 Technical Report
- GPT-4 Technical Report
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
- Best Practices for Text Annotation with Large Language Models
- Simulating Social Media Using Large Language Models to Evaluate Alternative News Feed Algorithms
- Qwen3 Technical Report
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution
- Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
- BotSim: LLM-Powered Malicious Social Botnet Simulation
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering