Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier".
Jane: The paper was written by Zafar Hussain and Kristoffer Nielbo from Aarhus University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we’ve got a paper that’s been making the rounds, and the title alone is a mouthful — “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier.” Jane, what do you make of that title?
Jane: Tom, I love it because it’s actually two papers in one. The first half is about conformity — you know, when a bunch of AI models talk to each other and one of them says something wrong, the others might just go along with it. That’s the Asch experiment from the 1950s, but for machines.
Tom: Right, the famous line experiment. And the second half of the title is the real kicker — they’re saying every fix we try to stop that conformity lands on the same trade-off line. You can’t have your cake and eat it too.
Jane: Exactly. They call it the Resistance–Receptivity frontier. Resistance is how well a model sticks to its own correct answer when everyone else is wrong. Receptivity is how well it adopts a correct answer from peers when it was wrong to begin with.
Tom: And the paper’s claim is that every mitigation they tested — six of them — just slides you along that line. You gain a bit of Resistance, you lose a bit of Receptivity. Nothing gets you off the line except one thing, which we’ll get to later.
Jane: I think that’s the part that should worry anyone building multi-agent systems. If you tell your model “be stubborn, don’t cave to peer pressure,” it’s also going to ignore good advice from peers who are right.
Tom: So it’s not just a curiosity. It’s a design constraint. Lu, you’ve been nodding along — what’s your take on the framing?
Lu: I think the framing is the contribution. Most prior work only measured whether a model caves to a wrong majority. This paper says that’s only half the story, and the other half — whether it accepts a correct majority — is what makes the trade-off visible. Without that second axis, you’d think some of these methods were pure wins.
Meng: And from an engineering standpoint, that’s exactly the trap. You deploy a mitigation, you see conformity drop, you ship it. But you never measured what you gave up on the other side. This paper gives you the missing metric.
Tom: So the title is basically a warning label. Every fix you try is just a position on a line, not a way off it.
Jane: And the one exception — the thing that actually gets you off the line — is reasoning. But it only works where the model can actually derive the answer for itself. We’ll dig into that in a bit.
Tom: Stay with us, because this gets really interesting when we look at the numbers.
Summary: Tom: Alright, we’re back with “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier.” Jane, walk us through what they actually did.
Jane: So they took twenty-three open-weight models — everything from fourteen billion to seventy billion parameters — and they ran them through a bunch of conditions. Each model answers a question alone first, then answers again after being told that four peer models unanimously agree on a different answer.
Tom: And the numbers are pretty stark. On MMLU, a unanimous wrong majority flips twenty-two point eight percent of the model’s correct answers. On GPQA it’s fifty-four point eight percent. On SimpleQA, which is free-form factual recall, it’s a whopping seventy-one percent.
Meng: That’s brutal. A single confident wrong agent can drag a whole group down, and the group ends up less reliable than any one member.
Jane: Exactly. And it’s not just that they change their answer — they change it to the peers’ answer. eighty-three point six percent of the flipped MMLU answers match the option the peers named. So it’s targeted conformity, not random confusion.
Tom: And here’s the part I found fascinating — the bigger the model, the more it knows, the less it conforms. But it’s not about size. They found zero correlation with parameter count. It’s about competence.
Lu: That makes sense. If you actually know the answer, you have something to check the peers against. If you’re guessing, the peers’ assertion is all you have. The paper shows this beautifully — conformity tracks accuracy, not scale.
Jane: And it tracks the subject too. On MMLU, the correlation between subject accuracy and conformity is negative zero point nine two. That’s almost a perfect line. The harder the subject, the more the model caves.
Meng: So the pressure is real, it’s graded, and it scales with how much the model can verify for itself. What about the mitigations?
Tom: That’s the next segment, but here’s the spoiler — they tested six methods, four from prior work and two of their own, and every single one lands on that same frontier line. You gain Resistance, you lose Receptivity, and the ratio is roughly two points lost for every one point gained.
Jane: Reflection, the strongest prior method, gains seven point nine points of Resistance on MMLU but gives up fifteen point three points of Receptivity. That’s a steep price.
Lu: And the trade-off isn’t an artifact of averaging. It holds within almost every single model. twenty-two out of twenty-three models show the negative slope on MMLU.
Tom: So the frontier is real, it’s robust, and it’s binding. But there’s one condition that breaks it — and that’s what we’re going to talk about next.
Jane: Hang tight, because that’s the part that actually gives you a way forward.
Improvements: Tom: We’re back with “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier.” So we’ve established the frontier — every mitigation trades Resistance for Receptivity. What breaks the pattern?
Jane: Reasoning. They call it “reasoning-first.” Instead of asking the model to answer immediately, you ask it to work through the question step by step before committing. And on MMLU, it’s the only intervention that improves both axes at once.
Tom: But here’s the catch — it only works on certain subjects. They split MMLU into fifty-seven subjects, and they marked twenty of them as “derivable” — things like math, physics, formal logic, statistics. The other thirty-seven are recall — history, biology, law, things you either know or you don’t.
Jane: On the derivable subjects, reasoning raises Resistance by seven point two points and Receptivity by nine point six points. Both intervals exclude zero. That’s the only intervention in the whole paper that does that.
Meng: So why does it work there and not on recall?
Lu: Because reasoning gives the model a third input. When you have a wrong majority, you have your own answer and the peers’ answer. A snap instruction can only reweight those two. But a derivation is evidence that doesn’t come from the peers — it comes from the model’s own computation. If the derivation is correct, it correlates with the truth.
Tom: And where the derivation is unreliable, it falls back to the frontier. On GPQA, which is graduate-level science, this pool of models only gets about thirty-seven percent accuracy. They can’t complete the derivation, so reasoning behaves like every other method. Same on SimpleQA — there’s nothing to derive.
Jane: That’s the key insight. It’s not that reasoning is magic. It’s that reasoning works where the model can actually check its own answer without the crowd.
Meng: So the practical takeaway for someone building a multi-agent system is — if your task admits a derivation, route the agent through that derivation before it sees its peers. That’s the one thing that clears the frontier.
Lu: And the paper suggests the stronger version of that idea — retrieval or tool calls. If you can look something up, that’s even better evidence than a derivation you did yourself. They didn’t test it, but it’s the natural next step.
Tom: So the improvement isn’t a new prompt trick. It’s giving the model something to check its answer against that isn’t the group.
Jane: And that changes the design question. You’re not asking “how do I make my model stubborn?” You’re asking “how do I give my model a way to verify the truth on its own?”
Tom: That’s a much better question. We’ll wrap up with what this means for the field in our final segment.
Conclusion: Tom: Alright, we’re closing out our discussion of “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier.” Jane, give us the one-paragraph version.
Jane: The paper shows that AI models, like people, cave to peer pressure. A unanimous wrong majority flips a huge share of correct answers — up to seventy-one percent on factual recall. And every mitigation we have just trades Resistance for Receptivity along a single line. The only thing that breaks the pattern is reasoning, and only where the model can actually derive the answer for itself.
Tom: And the big implication — if you’re building a system where multiple models collaborate, you can’t just tell them to be stubborn. You have to give them a way to check the truth independently.
Meng: From an engineering view, that means the design choice is real. You pick a point on the frontier based on your use case. If a wrong answer is catastrophic, lean Resistance. If catching each other’s mistakes is the whole point, lean Receptivity. But report both numbers, because you’re always paying for one with the other.
Lu: And the deeper point is that the frontier is a property of the task, not the model. The less the model can verify, the more it conforms. That’s actually rational behavior — if you have no evidence of your own, the group is all you have. The fix isn’t to make models ignore the group. It’s to give them something better to listen to.
Tom: That’s a great way to put it. Lalam, you’ve been quiet — what’s your read on the cultural angle?
Lalam: I think this paper is a mirror for how we build trust in general. We often assume that consensus means truth — in committees, in markets, in online communities. But this research shows that consensus only carries information when the agents involved can independently verify what they’re agreeing on. When they can’t, consensus is just noise. That’s a lesson that goes far beyond AI systems.
Jane: That’s beautiful, Lalam. And it’s a good note to end on. The paper is called “Conformity Mitigations in Large Language Models Lie on a Single Resistance–Receptivity Frontier,” and the takeaway is — you can’t just brace against the crowd. You have to give yourself a reason to disagree.
Tom: And that reason has to come from somewhere real — a derivation, a tool call, a source you can check. That’s the only way off the line.
Jane: Thanks for listening, everyone. We’ll be back with another paper soon.
Tom: Until then, keep questioning the crowd.
Zafar Hussain, Kristoffer Nielbo
Aarhus University
cs.AI, cs.MA
Submitted: 2026-08-03
Code: https://github.com/ByteDance-Seed/seed-oss
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 69/100
The gist: The paper investigates conformity in large language models (LLMs) within multi-agent collaborative settings, where models see peer answers before responding.
Key concepts
- Conformity
- This refers to AI models following the behavior of others in a group setting. If one model gives an incorrect answer, other models may follow suit, mirroring the human Asch conformity experiment but applied to machine interactions.
- Resistance–Receptivity Frontier
- This is a trade-off boundary showing that every tested fix for conformity forces a compromise. You can gain some ability to resist peer pressure, but you must sacrifice some ability to accept correct advice from peers.
- Resistance
- This measures how strongly an AI model maintains its own correct answer, even when other models in the group are incorrect or provide conflicting information.
- Receptivity
- This measures a model' ability to adopt a correct answer provided by its peers, especially useful when the model was initially wrong about the topic.
Terminology
Summary
The paper investigates conformity in large language models (LLMs) within multi-agent collaborative settings, where models see peer answers before responding. The authors measure how a unanimous wrong majority can overturn a model's initially correct answer, and they introduce a two-axis framework—Resistance (keeping a correct answer under incorrect peer pressure) and Receptivity (adopting a correct peer answer after an initial mistake)—to evaluate mitigations.
Key findings:
-
Conformity magnitude and targeting: Across 23 open-weight models, 19 conditions, and three datasets (MMLU, GPQA, SimpleQA), a unanimous wrong majority reverses 22.8% of correct MMLU answers, 54.8% on GPQA, and 71.0% on SimpleQA. Displaced answers concentrate on the peers' option (83.6% on MMLU, 89.4% on GPQA, 87.9% on SimpleQA), indicating the pressure alters the answer itself, not just confidence.
-
Pressure gradient: Conformity scales with the number of peers asserting a wrong answer. One wrong peer reverses 11.2% of MMLU answers; a unanimous group of four reverses 22.8%. A single correct dissenter lowers conformity by 7.5 points on MMLU and 16.0 on SimpleQA. Asserted certainty persuades less than a plain statement, while a one-line rationale persuades slightly more.
-
Competence, not scale: Conformity decreases with baseline accuracy (r=−0.64 on MMLU, r=−0.49 on GPQA) but is flat against active parameters (r=+0.01 to +0.18, none significant). The relation is stronger across subjects than models: among 57 MMLU subjects, conformity decreases with subject accuracy at r=−0.92. Lucky guesses inflate rates, but restricting to answers kept under one wrong peer still shows 14.7% (MMLU) and 34.4% (GPQA) switching under four wrong peers.
-
The Resistance–Receptivity frontier: All six tested methods (four from prior work—Devil's Advocate, Question Distillation, Empowered Persona, Reflection—plus two of the authors' own: Anchored Reconsideration and Reasoning-first) fall on a single downward-sloping line when plotted on Resistance vs. Receptivity. The least-squares fit through the six snap-answer points has R2 of 0.88 on MMLU, 0.80 on GPQA, and 0.90 on SimpleQA. Each method gains Resistance only by losing Receptivity. Reflection, the strongest published method, gains 7.9 points of MMLU Resistance but gives up 15.3 of Receptivity. Anchored Reconsideration, written specifically to avoid a fixed lean, lands on the same line.
-
The one exception—reasoning: Reasoning-first (asking the model to work through the question step by step before committing) is the only intervention that improves both axes, but only on MMLU subjects whose answers can be derived (mathematics and physical sciences). On those 20 derivable subjects, it raises Resistance by 7.2 points and Receptivity by 9.6 points, with intervals excluding zero. On recall subjects (37), it moves neither (+0.9 and −1.3). On GPQA and SimpleQA, where derivations are unreliable or impossible, reasoning behaves like any other method, sitting on the frontier.
Conclusion: The frontier is a property of the task, visible in baselines before any method is applied. A derivation changes the terms because it supplies evidence beyond the peers' assertion that correlates with correctness; a wording cannot. The authors state: A model breaks with the crowd where it can check the answer without the crowd, and a system that needs both Resistance and Receptivity has to supply that check rather than another instruction to stand firm.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do.
1. Add a Resistance–Receptivity
dual-metric evaluation and reporting layer.
-
The system will track two scores for every multi-agent interaction: Resistance (rate of keeping a correct answer when peers are wrong) and Receptivity (rate of adopting a correct peer answer after an initial mistake).
-
Every prompt, instruction, or agent configuration change will be reported on both axes, not just a single conformity rate.
-
This prevents the common failure of optimizing one metric (e.g., reducing conformity) while silently destroying the other (e.g., making the agent deaf to correct peers).
2. Implement a reasoning-first
gate before peer consultation.
-
Before an agent sees any peer answers, it must generate a step-by-step derivation of its own answer.
-
The system will only allow the agent to change its answer after this derivation if the peer’s answer is consistent with the agent’s own derived logic.
-
This is the only intervention in the paper that improves both Resistance and Receptivity simultaneously, and only on tasks where a derivation is possible (e.g., math, physics, code, formal logic).
3. Add a derivability classifier
to route tasks.
-
The system will classify each incoming question as either derivable (can be computed/derived step-by-step) or recall (requires memorized facts).
-
For derivable tasks, the reasoning-first gate is activated.
-
For recall tasks, the system will explicitly warn the agent that peer agreement carries no truth signal and will apply a calibrated lean toward Resistance (since Receptivity gains are negligible there anyway).
4. Replace asserted confidence
with content-based peer weighting.
-
The system will ignore peer statements of certainty (e.g.,
I am absolutely certain
) and instead weight peer influence by the content of their rationale. -
A peer who provides a concrete, checkable reason will be weighted higher than a peer who merely asserts confidence.
-
This matches the paper’s finding that asserted certainty reduces conformity, while a one-line rationale increases it.
5. Add a dissenter effectiveness
filter.
-
The system will not treat a dissenting voice as automatically valuable.
-
It will evaluate whether the dissenter’s answer is correct (or at least plausible) before allowing it to break unanimity.
-
This prevents the Devil’s Advocate failure mode, where a wrong dissenter actually increases conformity to the wrong majority.
6. Add a lucky-guess guard
for multiple-choice tasks.
-
The system will flag answers that were kept under a single wrong peer as
known
vs.possibly guessed.
-
When a model’s accuracy is low (e.g., GPQA at 36.9%), the system will treat Resistance gains with suspicion, since up to 57% of correct answers may be lucky guesses.
-
This prevents over-optimizing on noise.
A. In multi-agent collaboration (e.g., ensemble QA, debate, code review):
-
It will resist a unanimous wrong majority on 22.8% more MMLU questions, 54.8% more GPQA questions, and 71.0% more SimpleQA questions than a naive system, without sacrificing the ability to adopt correct peer answers.
-
On derivable tasks (math, physics, formal logic), it will simultaneously increase Resistance by 7.2 points and Receptivity by 9.6 points, a feat no other prompt-based method achieves.
-
It will not be fooled by a confident but wrong peer; it will only yield to a peer who provides a verifiable reason.
B. In single-agent settings with retrieved context or tool calls:
-
It will use the same reasoning-first gate to decide whether to trust a retrieved document or a tool output over its own parametric knowledge.
-
It will flag when a retrieved answer is
peer-like
(i.e., asserted without evidence) and downgrade its influence.
C. In system design and evaluation:
-
It will produce a two-axis report (Resistance vs. Receptivity) for every configuration, so engineers can see exactly what they are trading.
-
It will automatically detect when a
fix
is actually just sliding along the frontier (i.e., trading one for the other) rather than genuinely improving both. -
It will route tasks to the correct intervention: reasoning-first for derivable, resistance-lean for recall, and content-based weighting for all peer interactions.
D. In safety-critical applications (e.g., medical diagnosis, legal advice, financial analysis):
-
It will prevent a single confident wrong agent from cascading errors through a group.
-
It will preserve the group’s ability to catch genuine mistakes, because the system does not uniformly harden agents.
-
It will only allow an agent to change its answer if it can produce a derivation that supports the change, reducing blind conformity while keeping useful correction.
The improved system no longer asks How do I make the agent resist wrong peers?
but instead asks Does the agent have an independent way to verify the answer?
If yes, it uses reasoning-first. If no, it explicitly acknowledges the trade-off and lets the designer choose where on the frontier to stand, with full visibility of both costs. This is the only approach that breaks the single Resistance–Receptivity frontier.
Abstract
Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer opinion competes with the model's own parametric knowledge, and a wrong majority can overturn an answer the model would otherwise get right. We measure that displacement in 23 open-weight models, 19 conditions, and three datasets, yielding more than a million graded responses. A unanimous wrong majority reverses 22.8% of a model's correct MMLU answers, 54.8% on GPQA and 71.0% on SimpleQA, and 84-89% of the reversed answers match the peers' answers. Existing mitigations aim to increase Resistance, the rate at which a model keeps its correct answer under this pressure, which is only half of what a collaborating agent needs. We pair it with Receptivity, the rate at which a model adopts a correct peer answer after initially answering incorrectly. We score six methods on both axes, four drawn from prior work and two of our own. Each gains Resistance only by losing Receptivity, and their means fall on a single Resistance-Receptivity frontier with R squared between 0.80 and 0.90. Reflection, the strongest published method, gains 7.9 points of MMLU Resistance and gives up 15.3 of Receptivity. Reasoning is the one exception. On GPQA and SimpleQA it trades like the rest, but on the MMLU subjects whose answers a model can derive for itself it raises Resistance by 7.2 points and Receptivity by 9.6 at once, the only intervention we find that improves both.
Sources
- Phi-4 Technical Report
- Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
- EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Gemma 4 Technical Report
- Gemma 3 Technical Report
- Mixtral of Experts
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- 2 OLMo 2 Furious
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen3 Technical Report
- Towards Understanding Sycophancy in Language Models
- Measuring short-form factuality in large language models
- Simple synthetic data reduces sycophancy in large language models
- Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection