Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier

summary

Video file (mp4)

The gist

The paper investigates conformity in large language models (LLMs) within multi-agent collaborative settings, where models see peer answers before responding.

In short

The episode discusses research on LLM conformity, where models tend to follow a unanimous wrong majority. The paper establishes that all current fixes are limited by a trade-off between Resistance and Receptivity. The hosts conclude that the only way to break this limitation is through reasoning, which allows the model to independently verify its own answer.

Key concepts

Conformity
This refers to AI models following the behavior of others in a group setting. If one model gives an incorrect answer, other models may follow suit, mirroring the human Asch conformity experiment but applied to machine interactions.
Resistance–Receptivity Frontier
This is a trade-off boundary showing that every tested fix for conformity forces a compromise. You can gain some ability to resist peer pressure, but you must sacrifice some ability to accept correct advice from peers.
Resistance
This measures how strongly an AI model maintains its own correct answer, even when other models in the group are incorrect or provide conflicting information.
Receptivity
This measures a model' ability to adopt a correct answer provided by its peers, especially useful when the model was initially wrong about the topic.

Terminology used across episodes

This episode discusses

The paper

Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier · Read on arXiv

Zafar Hussain, Kristoffer Nielbo

Aarhus University

Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer opinion competes with the model's own parametric knowledge, and a wrong majority can overturn an answer the model would otherwise get right. We measure that displacement in 23 open-weight models, 19 conditions, and three datasets, yielding more than a million graded responses. A unanimous wrong majority reverses 22.8% of a model's correct MMLU answers, 54.8% on GPQA and 71.0% on SimpleQA, and 84-89% of the reversed answers match the peers' answers. Existing mitigations aim to increase Resistance, the rate at which a model keeps its correct answer under this pressure, which is only half of what a collaborating agent needs. We pair it with Receptivity, the rate at which a model adopts a correct peer answer after initially answering incorrectly. We score six methods on both axes, four drawn from prior work and two of our own. Each gains Resistance only by losing Receptivity, and their means fall on a single Resistance-Receptivity frontier with R squared between 0.80 and 0.90. Reflection, the strongest published method, gains 7.9 points of MMLU Resistance and gives up 15.3 of Receptivity. Reasoning is the one exception. On GPQA and SimpleQA it trades like the rest, but on the MMLU subjects whose answers a model can derive for itself it raises Resistance by 7.2 points and Receptivity by 9.6 at once, the only intervention we find that improves both.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier".

Jane: The paper was written by Zafar Hussain and Kristoffer Nielbo from Aarhus University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we’ve got a paper that’s been making the rounds, and the title alone is a mouthful — “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier.” Jane, what do you make of that title?

Jane: Tom, I love it because it’s actually two papers in one. The first half is about conformity — you know, when a bunch of AI models talk to each other and one of them says something wrong, the others might just go along with it. That’s the Asch experiment from the 1950s, but for machines.

Tom: Right, the famous line experiment. And the second half of the title is the real kicker — they’re saying every fix we try to stop that conformity lands on the same trade-off line. You can’t have your cake and eat it too.

Jane: Exactly. They call it the Resistance–Receptivity frontier. Resistance is how well a model sticks to its own correct answer when everyone else is wrong. Receptivity is how well it adopts a correct answer from peers when it was wrong to begin with.

Tom: And the paper’s claim is that every mitigation they tested — six of them — just slides you along that line. You gain a bit of Resistance, you lose a bit of Receptivity. Nothing gets you off the line except one thing, which we’ll get to later.

Jane: I think that’s the part that should worry anyone building multi-agent systems. If you tell your model “be stubborn, don’t cave to peer pressure,” it’s also going to ignore good advice from peers who are right.

Tom: So it’s not just a curiosity. It’s a design constraint. Lu, you’ve been nodding along — what’s your take on the framing?

Lu: I think the framing is the contribution. Most prior work only measured whether a model caves to a wrong majority. This paper says that’s only half the story, and the other half — whether it accepts a correct majority — is what makes the trade-off visible. Without that second axis, you’d think some of these methods were pure wins.

Meng: And from an engineering standpoint, that’s exactly the trap. You deploy a mitigation, you see conformity drop, you ship it. But you never measured what you gave up on the other side. This paper gives you the missing metric.

Tom: So the title is basically a warning label. Every fix you try is just a position on a line, not a way off it.

Jane: And the one exception — the thing that actually gets you off the line — is reasoning. But it only works where the model can actually derive the answer for itself. We’ll dig into that in a bit.

Tom: Stay with us, because this gets really interesting when we look at the numbers.

Summary: Tom: Alright, we’re back with “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier.” Jane, walk us through what they actually did.

Jane: So they took twenty-three open-weight models — everything from fourteen billion to seventy billion parameters — and they ran them through a bunch of conditions. Each model answers a question alone first, then answers again after being told that four peer models unanimously agree on a different answer.

Tom: And the numbers are pretty stark. On MMLU, a unanimous wrong majority flips twenty-two point eight percent of the model’s correct answers. On GPQA it’s fifty-four point eight percent. On SimpleQA, which is free-form factual recall, it’s a whopping seventy-one percent.

Meng: That’s brutal. A single confident wrong agent can drag a whole group down, and the group ends up less reliable than any one member.

Jane: Exactly. And it’s not just that they change their answer — they change it to the peers’ answer. eighty-three point six percent of the flipped MMLU answers match the option the peers named. So it’s targeted conformity, not random confusion.

Tom: And here’s the part I found fascinating — the bigger the model, the more it knows, the less it conforms. But it’s not about size. They found zero correlation with parameter count. It’s about competence.

Lu: That makes sense. If you actually know the answer, you have something to check the peers against. If you’re guessing, the peers’ assertion is all you have. The paper shows this beautifully — conformity tracks accuracy, not scale.

Jane: And it tracks the subject too. On MMLU, the correlation between subject accuracy and conformity is negative zero point nine two. That’s almost a perfect line. The harder the subject, the more the model caves.

Meng: So the pressure is real, it’s graded, and it scales with how much the model can verify for itself. What about the mitigations?

Tom: That’s the next segment, but here’s the spoiler — they tested six methods, four from prior work and two of their own, and every single one lands on that same frontier line. You gain Resistance, you lose Receptivity, and the ratio is roughly two points lost for every one point gained.

Jane: Reflection, the strongest prior method, gains seven point nine points of Resistance on MMLU but gives up fifteen point three points of Receptivity. That’s a steep price.

Lu: And the trade-off isn’t an artifact of averaging. It holds within almost every single model. twenty-two out of twenty-three models show the negative slope on MMLU.

Tom: So the frontier is real, it’s robust, and it’s binding. But there’s one condition that breaks it — and that’s what we’re going to talk about next.

Jane: Hang tight, because that’s the part that actually gives you a way forward.

Improvements: Tom: We’re back with “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier.” So we’ve established the frontier — every mitigation trades Resistance for Receptivity. What breaks the pattern?

Jane: Reasoning. They call it “reasoning-first.” Instead of asking the model to answer immediately, you ask it to work through the question step by step before committing. And on MMLU, it’s the only intervention that improves both axes at once.

Tom: But here’s the catch — it only works on certain subjects. They split MMLU into fifty-seven subjects, and they marked twenty of them as “derivable” — things like math, physics, formal logic, statistics. The other thirty-seven are recall — history, biology, law, things you either know or you don’t.

Jane: On the derivable subjects, reasoning raises Resistance by seven point two points and Receptivity by nine point six points. Both intervals exclude zero. That’s the only intervention in the whole paper that does that.

Meng: So why does it work there and not on recall?

Lu: Because reasoning gives the model a third input. When you have a wrong majority, you have your own answer and the peers’ answer. A snap instruction can only reweight those two. But a derivation is evidence that doesn’t come from the peers — it comes from the model’s own computation. If the derivation is correct, it correlates with the truth.

Tom: And where the derivation is unreliable, it falls back to the frontier. On GPQA, which is graduate-level science, this pool of models only gets about thirty-seven percent accuracy. They can’t complete the derivation, so reasoning behaves like every other method. Same on SimpleQA — there’s nothing to derive.

Jane: That’s the key insight. It’s not that reasoning is magic. It’s that reasoning works where the model can actually check its own answer without the crowd.

Meng: So the practical takeaway for someone building a multi-agent system is — if your task admits a derivation, route the agent through that derivation before it sees its peers. That’s the one thing that clears the frontier.

Lu: And the paper suggests the stronger version of that idea — retrieval or tool calls. If you can look something up, that’s even better evidence than a derivation you did yourself. They didn’t test it, but it’s the natural next step.

Tom: So the improvement isn’t a new prompt trick. It’s giving the model something to check its answer against that isn’t the group.

Jane: And that changes the design question. You’re not asking “how do I make my model stubborn?” You’re asking “how do I give my model a way to verify the truth on its own?”

Tom: That’s a much better question. We’ll wrap up with what this means for the field in our final segment.

Conclusion: Tom: Alright, we’re closing out our discussion of “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier.” Jane, give us the one-paragraph version.

Jane: The paper shows that AI models, like people, cave to peer pressure. A unanimous wrong majority flips a huge share of correct answers — up to seventy-one percent on factual recall. And every mitigation we have just trades Resistance for Receptivity along a single line. The only thing that breaks the pattern is reasoning, and only where the model can actually derive the answer for itself.

Tom: And the big implication — if you’re building a system where multiple models collaborate, you can’t just tell them to be stubborn. You have to give them a way to check the truth independently.

Meng: From an engineering view, that means the design choice is real. You pick a point on the frontier based on your use case. If a wrong answer is catastrophic, lean Resistance. If catching each other’s mistakes is the whole point, lean Receptivity. But report both numbers, because you’re always paying for one with the other.

Lu: And the deeper point is that the frontier is a property of the task, not the model. The less the model can verify, the more it conforms. That’s actually rational behavior — if you have no evidence of your own, the group is all you have. The fix isn’t to make models ignore the group. It’s to give them something better to listen to.

Tom: That’s a great way to put it. Lalam, you’ve been quiet — what’s your read on the cultural angle?

Lalam: I think this paper is a mirror for how we build trust in general. We often assume that consensus means truth — in committees, in markets, in online communities. But this research shows that consensus only carries information when the agents involved can independently verify what they’re agreeing on. When they can’t, consensus is just noise. That’s a lesson that goes far beyond AI systems.

Jane: That’s beautiful, Lalam. And it’s a good note to end on. The paper is called “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier,” and the takeaway is — you can’t just brace against the crowd. You have to give yourself a reason to disagree.

Tom: And that reason has to come from somewhere real — a derivation, a tool call, a source you can check. That’s the only way off the line.

Jane: Thanks for listening, everyone. We’ll be back with another paper soon.

Tom: Until then, keep questioning the crowd.

More episodes

← Home