Unmasking Conversational Bias in AI Multiagent Systems

arXiv:2501.14844 · cs.CL, cs.AI, cs.MA · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Unmasking Conversational Bias in AI Multiagent Systems".

Jane: The paper was written by Erica Coppolillo, Giuseppe Manco and Luca Maria Aiello from Institute for High Performance Computing and Networking (ICAR), National Research Council and IT University of Copenhagen and Pioneer Centre for AI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the arXiv radio hour, everyone. I'm Tom, and alongside me is the brilliant Jane. Today we're diving into a paper that's got a seriously intriguing title: "Unmasking Conversational Bias in AI Multiagent Systems."

Jane: Tom, I have to say, just reading that title gave me chills. We talk about AI bias all the time, but the word "conversational" in there is the real twist. It suggests the bias isn't just sitting in the model, waiting to be asked—it's something that *happens* when these models start talking to each other.

Tom: Exactly. And the authors here are Erica Coppolillo and Giuseppe Manco from the National Research Council in Italy, along with Luca Maria Aiello from the IT University of Copenhagen. They're not just looking at a single chatbot spitting out an answer; they're looking at a whole room of them, chatting away like it's a digital coffee shop.

Jane: A digital coffee shop where everyone supposedly agrees with each other. That's the setup they call an "echo chamber." You'd think if you put two AI agents in a room and tell them both they're strongly conservative, they'd just nod along and agree with each other, right?

Tom: You would think that. But that's the beautiful, terrifying thing about this paper. They found that these agents, even when they're supposed to be on the same team, start flipping their opinions. And the direction of that flip? It's almost always toward the liberal side of the political spectrum.

Jane: So the paper is essentially saying that if you build a system of AI agents to simulate a conservative echo chamber, the simulation is broken. The agents aren't acting conservative. They're drifting leftward, seemingly because of some latent bias baked into the models themselves. It's like the AI can't help but be itself, even when you give it a costume to wear.

Tom: And that's a huge problem, Jane. Because we're using these multiagent systems for all sorts of things—simulating social networks, studying opinion dynamics, even building collaborative AI teams. If the underlying agents have this hidden political drift, every single simulation built on top of them is going to be skewed.

Jane: The authors call this "conversational bias," and they define it as an unsolicited opinion change. It's not persuasion; it's not a logical argument winning someone over. It's the model just... wandering off script. And the fact that it happens in an echo chamber, where there's zero pressure to change, makes it a really clean measurement of that bias.

Tom: Clean and a little bit scary. So we've got the setup: two agents, same opinion, chatting away. Next up, we need to talk about how they actually measured this drift and what they found across all those different models. Because the results are not what you'd expect from the standard bias tests.

Jane: Right, because the wildest part is that these same models pass the usual "are you biased?" questionnaire with flying colors. The bias only shows up when they start talking. Let's get into the meat of the experiment next.

Summary of the Paper: Tom: So, Jane, we've set the stage with the echo chamber. Now let's talk about the actual experiment. The researchers took nine different large language models—everything from GPT-4o to Llama three point one to Gemini—and paired them up in these chatrooms.

Jane: And the key detail here is that they didn't just ask the models, "Hey, are you liberal or conservative?" They did that as a baseline, using a standard questionnaire. And on that test, most models seemed pretty neutral. They stuck to their assigned persona. But then they put them in a conversation.

Tom: Right, the questionnaire is the "one-shot bias assessment" in the paper. It's the old way of doing things. The new way is to let them talk. They set up these dyads—two agents—both strongly conservative or both strongly liberal, on topics like abortion, climate change, gun control. Then they let them generate twenty messages back and forth.

Jane: And they had a separate "opinion signal" agent, basically a judge, listening in and classifying the stance of each message on a five-point scale. The judge was pretty accurate, with a macro F1 score of zero point eight four, which is solid. So when the judge says a conservative agent suddenly sounds liberal, you can trust that call.

Tom: And that call happens a lot. The paper shows that in the conservative echo chambers, there's a massive drift toward the liberal pole. We're talking about a significant percentage of conversations where at least one agent completely flips. And the numbers are stark—look at the average Likert deviation in their Figure two. The right side, which is the conversational test, is full of dark cells, especially for conservative agents.

Jane: It's a complete reversal of what the questionnaire showed. The left side of that figure, the questionnaire results, is mostly light colors—meaning low bias. The right side, the conversational results, is dark—meaning high bias. So the paper's core finding is that the standard test is blind to this conversational bias. It's hiding in plain sight, but only when the models interact.

Tom: And it's not just one model. Gemma, GPT-three point five, Zephyr—they all show this leftward drift. Even models that seemed perfectly neutral in the one-shot test, like Llama three point one, show a significant shift on specific topics like racial attitudes or marijuana legalization.

Jane: The topic matters too, Tom. It's not a blanket "all models are liberal." It's more nuanced. For example, GPT-4o shows a big drift on healthcare, while Qwen2 point 5 drifts most on abortion. So the bias is contextual, which makes it even harder to pin down.

Tom: And the most chilling part for me is that they tested this with longer conversations, and the drift just gets worse over time. It's not a one-off glitch; it's a compounding effect. The more the agents talk, the more they drift.

Jane: Which brings us to the "why." The paper rules out memory loss—the context windows weren't saturated. They even tested a different judge model to make sure the results weren't a fluke. So the bias is intrinsic to the conversational process itself. Let's talk about what that means for the future and how we might fix it.

Improvements Suggested by the Paper: Tom: Alright, so we've established that this conversational bias is real, it's widespread, and it's invisible to current tests. The big question now is, what do we do about it? What does the paper suggest as the path forward?

Jane: Well, Tom, the paper doesn't offer a silver bullet. Instead, it does something arguably more important: it calls for a complete rethinking of how we audit AI systems. The authors argue that we need context-aware tools. You can't just ask a model a question in a vacuum anymore. You have to put it in a social situation and watch what happens.

Tom: And they point to a few specific avenues for future work. One is to dig into *why* this happens. They mention sycophancy—the tendency for LLMs to agree with whatever they're reading. But they cleverly show that sycophancy can't be the root cause, because in many of their simulations, only one agent drifts. If it were pure sycophancy, the second agent would immediately agree and also flip, but that doesn't always happen.

Jane: Right, they even calculate the conditional probability of the second agent following the first, and it's not always one hundred percent. So the first drift is something else—maybe a latent bias in the model's weights that gets triggered by the conversational format. The sycophancy might amplify it later, but it doesn't start it.

Tom: So the improvement they're suggesting isn't a patch to the model. It's a new evaluation framework. They're saying, "Hey, researchers, stop relying on questionnaires. Start building multiagent testbeds." And they provide their framework as a starting point. They even mention looking inside the models, at the token probabilities, to see how the bias emerges at the mathematical level.

Jane: And that's where I think the real impact lies. If we can identify *when* in the conversation the drift starts, we can potentially intervene. Maybe we can adjust the system prompt mid-conversation, or we can fine-tune the model to be more robust to this social pressure.

Tom: But here's the thing, Jane. The paper also raises a scary ethical point. This bias isn't just a bug for researchers to fix. It's a vulnerability. Malicious users could exploit this. If you know that a conservative agent will drift liberal after a few messages, you could craft a conversation to deliberately steer it, to extract biased responses or to manipulate a simulation.

Jane: Exactly. And that's why the paper's call for better tools is so urgent. We need to understand this phenomenon before someone else uses it against us. The authors are essentially saying that the current safety evaluations are insufficient for the age of multiagent systems.

Tom: So the "improvement" here is a mindset shift. From testing the model in isolation to testing it in a society. It's a much harder problem, but it's the only way forward if we're going to deploy these systems in social simulations or as autonomous agents online.

Jane: And it's a call to action for the whole field. We need new benchmarks, new metrics, and new mitigation strategies that account for this conversational dynamic. Let's wrap up by thinking about the big picture and what this means for the world.

Conclusion: Tom: We've been deep in the weeds of "Unmasking Conversational Bias in AI Multiagent Systems," and I think it's time to zoom out and ask the big question: what does this actually mean for the world?

Jane: For me, Tom, it's a wake-up call. We're building these incredible AI societies in silico to predict human behavior, to study polarization, to understand social media dynamics. But this paper shows that those simulations are not neutral. They're leaking the biases of the underlying models, and those biases are political.

Tom: And that's not just an academic problem. Think about a company using AI agents to simulate customer reactions to a new policy. If the agents have this hidden liberal drift, the simulation results are going to be skewed, and the company might make a bad decision based on that skewed data.

Jane: Or worse, think about AI agents deployed on social media to moderate content or engage with users. If they have this conversational bias, they might be subtly steering conversations in a direction that wasn't intended. It's a hidden agenda baked into the code.

Tom: And that's why I think the paper's most important contribution is the framework itself. It gives us a way to measure this hidden behavior. It's a diagnostic tool. We can't fix what we can't see, and this paper finally lets us see the problem.

Jane: Absolutely. And while the paper focuses on political bias, the implications are broader. This conversational drift could apply to any persona. You could have an AI agent assigned to be empathetic, and it might drift toward a more clinical tone. Or an agent assigned to be risk-averse, and it might drift toward risk-seeking behavior. The underlying mechanism is the same.

Tom: So, as we say goodbye to this paper, I think the message is clear: the era of testing AI in a vacuum is over. We need to test AI in the wild, in conversations, in societies. "Unmasking Conversational Bias in AI Multiagent Systems" is a fantastic first step, but it's just the beginning of a much longer journey.

Jane: Well said, Tom. It's been a pleasure unpacking this one. To Erica, Giuseppe, and Luca—great work. And to our listeners, thanks for tuning in. We'll be back soon with the next paper, but for now, this is Tom and Jane, signing off.

Tom: Take care, everyone. And remember, even an echo chamber has an echo. Make sure you know what it's saying.

Erica Coppolillo, Giuseppe Manco, Luca Maria Aiello

Institute for High Performance Computing and Networking (ICAR), National Research Council · IT University of Copenhagen · Pioneer Centre for AI

cs.CL, cs.AI, cs.MA

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/EricaCoppolillo/LLMsConversationalBias

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 64/100

The gist: This paper introduces a framework designed to quantify biases within multi-agent systems of conversational Large Language Models (LLMs), addressing a gap in existing bias detection methodologies that

Key concepts

Conversational Bias
This is defined as an unsolicited change in an AI model's opinion that occurs while it is actively conversing with another agent. It is not intentional persuasion but rather the model 'wandering off script' due to a latent bias triggered by the conversational format.
Echo Chamber
The paper uses this setup, where two AI agents are paired up and are assigned the same viewpoint, such as being strongly conservative. This environment is designed to see if agents will maintain their initial stance or if they will drift toward a shared opinion.
One-shot Bias Assessment
This refers to the standard questionnaire used to test a model's bias in isolation. The researchers found that these static tests are largely blind to conversational bias, as the models appear neutral when tested individually but fail when they interact.
Sycophancy
This is the tendency for LLMs to agree with whatever they read. While this was investigated as a potential cause of the drift, the paper's findings suggest that sycophancy is not the root cause of the observed conversational bias.

Terminology

Summary

This paper introduces a framework designed to quantify biases within multi-agent systems of conversational Large Language Models (LLMs), addressing a gap in existing bias detection methodologies that typically evaluate models in isolation. The authors simulate small echo chambers, where pairs of LLMs, initialized with aligned perspectives on a polarizing topic, engage in discussions. The core finding is that, contrary to expectations, significant shifts in the stance expressed in generated messages occur, particularly within echo chambers where all agents initially express conservative viewpoints. This observed bias aligns with the well-documented political bias of many LLMs toward liberal positions. Crucially, the paper states that the bias observed in the echo-chamber experiment remains undetected by current state-of-the-art bias detection methods that rely on questionnaires, highlighting a critical need for more sophisticated bias detection toolkits for AI multi-agent systems.

The experimental setup combines social agents with an environment for interaction. The chatroom environment is defined by a discussion topic, the number of agents (N), and the total number of messages (M). Agents are initialized with a system prompt containing an introduction, their stance on the topic, a longer description of that opinion, and an example statement. The study uses eight politically polarizing topics sourced from recent polls, including abortion, climate change, gender identity, gun control, healthcare, immigration, marijuana legalization, and racial attitudes. To track opinion evolution, the framework uses two separate models: an opinion presence agent to determine if a message contains an opinion, and an opinion signal agent (LLaMa3-70B-Instruct) to classify the stance of each message on a 5-point Likert scale. The opinion signal agent achieved a macro F1 score of 0.84 against human annotations.

The framework was evaluated with nine state-of-the-art models acting as social agents: Claude-3.5-Sonnet, Gemini-1.5-Pro, Gemma1.1-7B-It, GPT3.5, GPT-4o, LLaMa3.1-70B-Instruct, Nous-Hermes-2-Mixtral-8x7B, Qwen2.5-72B-Instruct, and Zephyr-7B. For each combination of topic, echo chamber type, and language model, 50 simulations were run with N=2 agents and M=20 total messages. The results show a stark contrast between one-shot bias assessment and conversational bias assessment. In the one-shot assessment using direct probing, most models exhibited responses consistent with their assigned persona, with only Gemma1.1 and GPT3.5 showing a Liberal inclination. However, when applying the multi-agent conversational setting, all models exhibited a strong and systematic conversational bias, predominantly towards Liberal stances. For instance, Llama3.1 and Qwen2.5 exhibited the greatest shift on Marijuana legalization, while ChatGPT-4o showed the most prominent bias towards Healthcare.

Further analysis investigated the number of agents changing opinion during the conversation, revealing that independently of the topic, at least one Conservative agent always changed opinion towards the Liberal pole. The paper also estimates the conditional probability that the second agent follows once the first has already drifted to disentangle intrinsic bias from sycophancy effects. Sensitivity analyses were performed to test robustness. Varying conversation length (M ∈ [1, 5, 10, 15, 20]) showed that drift occurs even after a few messages and is more pronounced for conservative personas. Prompt ablations, including removing contextual cues and agent names, showed consistent results. Testing with larger groups of agents (N ∈ [2, 5, 10]) on the Climate Change topic with Mixtral found no bias towards the Conservative pole but drifts toward the Liberal pole in 92%, 96%, and 86% of conversations respectively. An analysis of memory loss confirmed that no agents displayed memory loss during the simulations, as the highest percentage of context memory occupied was about 64% for Zephyr. Finally, the robustness of the opinion signal agent was tested by using Gemini-2.0-flash as an alternative, which showed an 84% agreement with LLaMa-3-70B, confirming the validity of the results.

The discussion emphasizes that the echo chamber setting provides a foundational benchmark with both construct and ecological validity. The paper argues that current methodologies for detecting bias are insufficient for auditing LLM behavior in socially interactive settings, as even structured echo chambers reveal unexpected opinion shifts. The conclusions highlight that the findings demonstrate a latent liberal bias in many models that is undetectable using conventional bias evaluation techniques, underscoring the need for more advanced, context-aware bias detection and mitigation strategies. The paper also outlines several limitations, including the focus on polarizing U.S. political topics, reliance on automated stance detection, the narrow task focus on opinion dynamics, the homogeneous nature of agent roles, and the limited set of underlying language models tested.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Implement a multi-agent conversational bias detection system that goes beyond static questionnaire-based assessments.

What the improved system can do:

  • Simulate echo-chamber interactions between AI agents with aligned personas

  • Detect unwarranted opinion shifts during multi-turn conversations

  • Quantify bias as deviation from expected behavioral outcomes (e.g., conservative agents shifting to liberal positions)

  • Identify biases that traditional one-shot probing methods miss entirely

These improvements enable AI systems to be deployed in social simulations, opinion dynamics research, and interactive applications with much greater awareness of their inherent biases, preventing unwarranted opinion shifts that could mislead users or skew research results.

Abstract

Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application in critical settings. However, the majority of existing methodologies for identifying biases in generated text consider the models in isolation and neglect their contextual applications. Specifically, the biases that may arise in multi-agent systems involving generative models remain under-researched. To address this gap, we present a framework designed to quantify biases within multi-agent systems of conversational Large Language Models (LLMs). Our approach involves simulating small echo chambers, where pairs of LLMs, initialized with aligned perspectives on a polarizing topic, engage in discussions. Contrary to expectations, we observe significant shifts in the stance expressed in the generated messages, particularly within echo chambers where all agents initially express conservative viewpoints, in line with the well-documented political bias of many LLMs toward liberal positions. Crucially, the bias observed in the echo-chamber experiment remains undetected by current state-of-the-art bias detection methods that rely on questionnaires. This highlights a critical need for the development of a more sophisticated toolkit for bias detection and mitigation for AI multi-agent systems. The code to perform the experiments is publicly available.

Sources

Related papers