The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "The Collective Turing Test".
Jane: Large Language Models (LLMs) can generate social media conversations sufficiently realistic to deceive humans when reading them,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We started by looking at the title, "The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions," and it immediately tells us the paper is focused on a specific challenge: making AI discussions look genuinely human.
Jane: That title sets up the whole experiment perfectly; they aren't just testing one question or one response, but a whole group conversation to see if the overall flow feels authentic. It’s about simulating that messy, back-and-forth vibe we see online.
Lu: What’s interesting is that the authors are bringing together several researchers from different universities, which suggests this topic is becoming very important across multiple AI labs right now.
Meng: I noticed the authors list a few people from different areas; it tells me they're looking at this problem from various angles, maybe even considering how to solve it technically and ethically at the same time.
Lalam: It shows that simulating social interaction isn't just one thing; it involves language models, testing methods, and understanding how communities actually function together.
The paper's summary: Tom: Moving into the actual summary of "The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions," the authors explain they collected real human conversations from Reddit and then generated artificial ones using models like GPT-4o and Llama three 70B.
Jane: That’s a pretty straightforward setup; they took authentic posts from Reddit, got about twenty comments for each, and had the AI write new replies that matched the original comment's length. It’s essentially a controlled way to see if the generated text passes the human test.
Lu: The core finding they highlight is that Llama three 70B performed better than GPT-4o in mimicking those specific social media styles, suggesting that its training might have given it an edge in generating less polished, more organic responses.
Meng: That's a big piece of data for me because it implies that the way a model is trained can directly affect its perceived authenticity when dealing with informal content versus structured queries.
Lalam: It means we should pay close attention to what data we feed into models if our goal is to simulate real-world, casual human interactions instead of just formal ones.
The paper's improvements: Tom: Now for the part where the authors suggest how to make these simulations better, they point out that there are distinct cues people use to tell AI from human text, and those include tone, formality, and things like slang or profanity.
Jane: The paper found that textual style features were the main indicators participants used when distinguishing between the real Reddit conversation and the AI-generated one; eighty-two point five percent of participants mentioned at least one of these stylistic differences.
Lu: They also noted that conversation length has a non-linear effect, meaning success rates actually go up to a certain point—specifically length equals eight—and then start to drop off as the threads get much longer.
Meng: That length curve is really interesting from an engineering standpoint; it suggests that maintaining consistency over very long interactions is where the simulation starts breaking down for detection purposes.
Lalam: I think this gives us a clear roadmap for improving our systems: we need to focus on injecting emotional nuance and avoiding overly polite or perfectly structured language that we saw in the AI outputs.
Conclusion: Tom: So, wrapping up the paper "The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions," it’s clear that LLMs can certainly generate discussions that are mistaken for human-made, with a success rate of zero point six one overall.
Jane: The authors conclude that Llama three 70B is the better model for this task because it seems to capture the text of social media conversations more organically than GPT-4o, even though they found some limitations regarding emotional depth and personal storytelling in the AI outputs.
Lu: This result points toward a future where we can build simulations that are much closer to actual online community behavior, provided we keep iterating on how we condition those models.
Meng: Practically speaking, it means that for any application requiring simulated social media presence, choosing the right model and understanding the length dynamics will be essential for getting results you can actually use.
Lalam: Ultimately, this paper shows that while the simulation is convincing on a surface level, capturing true human nuance and emotional charge remains a challenge we need to tackle next.
Azza Bouleimen, Giordano De Marzo, Taehee Kim, `Nicolo Pagan` `Hannah Metzler`, `Silvia Giordano`, `Aniko Hannak´ ak´, and David Garcia
University of Zurich · University of Konstanz · Complexity Science Hub · University of Applied Sciences and Arts of Southern Switzerland
cs.CL, cs.AI, cs.SI
Submitted: 2025-10-29
Updated: 2025-10-29
Code: https://github.com/azza-bouleimen/turing-test-aiconversations11
Importance score: 76/100
The gist: Large Language Models (LLMs) can generate social media conversations sufficiently realistic to deceive humans when reading them, highlighting both a promising potential for social simulation and a
Key concepts
- Collective Turing Test
- This method involves showing participants side-by-side conversations: one real human discussion and one generated by an AI. Participants judge which is which. It tests the ability of an LLM to mimic natural, multi-user social media dialogue effectively.
- Model Selection Effect
- The choice of LLM significantly impacted performance; Llama 3 70B produced more human-like conversations than GPT-4o. This suggests that different models have different strengths in generating organic versus polished text based on their training data.
- Qualitative Cues for Detection
- Participants primarily detected AI conversations by analyzing textual style features, such as tone, formality, and the lack of slang or profanity. These stylistic differences are the main indicators used to distinguish machine-generated text from human writing.
- Stimuli Generation Strategy
- Researchers created 574 unique conversation examples by using minimal prompting to instruct LLMs to act like Reddit users. They varied three temperature settings and six comment lengths to create diverse stimuli for testing.
Terminology
Summary
Large Language Models (LLMs) can generate social media conversations sufficiently realistic to deceive humans when reading them, highlighting both a promising potential for social simulation and a warning message about the potential misuse of LLMs to generate new inauthentic social media content.
Key Findings
-
Overall, participants’ success rate was 0.61, meaning LLMs generated discussions that were mistaken to be human-made in 39% of the cases, surpassing the Turing Test threshold of 30%.
-
Model selection significantly affected performance:
Llama 3 70B outperformed GPT-4o in producing more human-like conversations.
This suggests Llama 3, due to its training data potentially including social media conversations, may generatemore organic and less polished responses than GPT-4o.
-
Conversation length showed a non-linear effect:
success rates increase up to a certain point (in this case, length = 8), beyond which success rates decline,
suggesting that very long threads may make it harder to maintain focus. -
Qualitative analysis revealed key cues for detection:
Textual style features served as the primary indicator participants used when distinguishing human from AI conversations, with 82.5% of responses mentioning at least one feature from this category.
These included differences in tone, text formality, and the absence of slang or profanity.
Methodology and Stimuli Generation
The study employed a Collective Turing Test
approach by presenting participants with side-by-side conversations: one authentic human conversation from Reddit and one artificially generated by an LLM (GPT-4o or Llama 3 70B). The stimuli were collected from the Pushshift dataset, specifically selecting 16 posts spanning 8 different topics: food, leisure, opinions, relationships, technology, movies, health, and sports.
For each post selected from Reddit (up to December 2023), the researchers scraped the first 20 comments. The conversational structure—which user replied to which other user under which thread—and the word counts for each comment were preserved while human comments were replaced with AI-generated ones.
The artificial conversations were generated using a minimal prompting strategy, instructing the models to behave like Reddit users without assigning any specific personas.
The core prompt directed the model to Write a reply with about [word count] words to comment ID [in reply to id],
tasking it to generate text of approximately the same length as the corresponding human comment. To ensure variation, conversations were generated using three different temperatures (0.2, 0.7, and 1.2) and presented at six different lengths (1, 2, 4, 6, 8, or 16 comments). This resulted in 574 different stimuli combining posts with AI versus human conversations for presentation to participants.
Participant Assessment and Analysis
Participants were asked to identify which conversation was human-made and to provide a qualitative assessment of how they distinguished the two. To ensure reliability, the annotation task involved six annotations per participant, followed by an additional attention test. A post-survey questionnaire was used to gather qualitative insights into the distinguishing features. The researchers developed a codebook for classifying participant responses based on three main categories: Format (visual cues), Style (tone, voice), and Content (explanation, authenticity). Inter-annotator agreement for these categories reached Cohen’s Kappa scores between 0.66 and 0.75 at both category and subcategory levels.
Model Comparison and Limitations
The analysis focused on the interaction between the LLM model used (GPT-4o vs. Llama 3 70B) and the conversation length on participant success rates, as well as qualitative cues. The statistical model indicated a significant effect of the LLM used to generate AI conversations,
with Llama 3 70B showing a lower success rate than GPT-4o. However, the study also found that conversation length generally has no significant effect on success rates (p-value = 0.36)
when treated as a continuous variable, though a categorical analysis showed an inverse-U relationship,
peaking at length = 8 before declining.
Qualitative findings highlighted limitations: participants noted that AI conversations were often overly polite, formal, and devoid of slang or profanity,
and they lacked authenticity, emotional nuance, and personal storytelling
compared to real social media conversations. This suggests current LLMs struggle to replicate toxic or emotionally charged behavior
central to studying online conflict.
Conclusion
The study concludes that LLMs can convincingly simulate human discourse on social media, with Llama 3 70B being superior for this task. However, limitations persist in replicating emotionally charged discourse, personal storytelling, or polarization,
indicating that careful prompt engineering and model selection are necessary for more realistic simulations.
Improvements for AI systems
As a fastidious and diligent AI researcher, I have analyzed this paper, The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions,
and identified several concrete avenues for improvement in AI systems, particularly those intended for social simulation or community interaction.
Here are the specific improvements I recommend:
)
-
Improve LLM Performance via Model Selection Strategy:
-
Enhance Conversational Realism through Prompt Engineering & Conditioning:
-
Develop Robust Validation and Benchmarking Protocols:
-
Implement Dynamic Conversation Length Adaptation for Detection Evasion:
)
- Improve LLM Performance via Model Selection Strategy:
The study found that Llama 3 70B outperformed GPT-4o in generating more human-like social media conversations (66% success rate vs. 56% for GPT-4o).
The improved AI system should implement a Model Steering Layer
that dynamically selects the underlying LLM based on the target simulation's requirements:
-
If the goal is to simulate highly informal, slang-heavy, or low-polish discourse (e.g., early Reddit threads), prioritize models like Llama 3 70B.
-
If the goal is high reliability, safety, or structured Q&A in a simulated environment, prioritize models like GPT-4o.
- Enhance Conversational Realism through Prompt Engineering & Conditioning:
The qualitative analysis showed that LLMs struggle with key human traits: lack of emotional nuance, absence of slang/profanity (often overly polite), and conformity (over-agreeing).
The improved AI system needs a sophisticated, layered prompt engineering strategy incorporating learned behavioral constraints:
-
Implement
Style Injection
modules that explicitly condition the model on specific stylistic features identified as human cues (e.g.,Inject subtle sarcasm,
Use common Reddit slang sparingly,
orAvoid excessive politeness
). -
Integrate a constraint mechanism against undesirable traits, such as a penalty for high conformity scores or a requirement to introduce low-frequency, contextually appropriate emotional language (like mild irony or frustration) to increase authenticity.
- Develop Robust Validation and Benchmarking Protocols:
The study highlighted the need for novel validation beyond simple Turing tests, noting that current methods often fail to capture operational validity in complex social dynamics.
The improved system must be validated using a multi-faceted Collective Turing Test
framework:
-
Move beyond binary identification (AI vs. Human) to a multi-dimensional assessment based on the participant cues identified (Format, Style, Content). The system should be benchmarked not just on overall success rate, but on its ability to mimic specific high-value features like
authenticity
andcontroversy.
-
Incorporate dynamic testing where the system is tested against diverse conversation lengths (as shown in Figure 2), specifically targeting the non-linear detection curve (peaking at length = 8) to understand where human fatigue or LLM inconsistency begins to dominate.
- Implement Dynamic Conversation Length Adaptation for Detection Evasion:
The research noted a non-linear relationship with length, peaking at 8 comments before declining, suggesting that very long threads (length > 16) may make detection harder or lead to participant fatigue.
The improved system should incorporate a Length Modulation Module
that dynamically adjusts its output structure based on the simulated conversation's context:
-
For simulations requiring high realism in short interactions (e.g., quick replies), the system should default to shorter, more tightly focused outputs.
-
For long-form discussions where sustained engagement is expected, the module should strategically introduce variations in response length and complexity (mimicking natural human ebb and flow) rather than maintaining a uniform word count across all turns. This counters the LLM tendency toward consistent verbosity when constrained by prompt length.
Sources
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- AI agents can coordinate beyond human scale
- Disinformation and Social Bot Operations in the Run Up to the 2017 French Presidential Election
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- Emergence of Scale-Free Networks in Social Interactions among Large Language Models
- Simulating Social Media Using Large Language Models to Evaluate Alternative News Feed Algorithms
- S$^3$: Social-network Simulation System with Large Language Model-Empowered Agents
- OASIS: Open Agent Social Interaction Simulations with One Million Agents
- Y Social: an LLM-powered Social Media Digital Twin
- VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models
- Human Variability vs. Machine Consistency: A Linguistic Analysis of Texts Generated by Humans and Large Language Models
- How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering