LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning".
Jane: The paper was written by Ning Liu from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we’re looking at a paper that’s got a wonderfully simple title: “LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning.” Jane, I have to say, that title alone tells you most of the story.
Jane: It really does, Tom. The idea is that instead of having one model check its own work, you get several different models to solve the problem independently, and then you take a vote. If they all agree on an answer, that’s a strong signal it’s correct. The paper calls this an “LLM-jury,” which is such a vivid way to think about it.
Tom: And the punchline is that this simple voting scheme beats the fancy trained reward models that people have been building. Those are the process reward models, or PRMs, that score each step of a solution. The paper shows that on competition math, the jury closes the entire gap to an oracle selector, while a model scoring its own candidates captures almost none of that gap.
Jane: Right, and that’s the part I find so striking. The paper says a model scoring its own twelve candidates captures essentially zero percent of the achievable selection gain over self-consistency. But three models that never see each other’s work capture one hundred percent of it. So the signal that separates a correct answer from a wrong one lives between models, not inside any single one.
Tom: That’s a profound observation, Jane. It’s like the classic idea that diverse perspectives catch errors that a single perspective misses. The paper formalizes this with something they call error decorrelation. When models are independently trained, their mistakes tend to be different, so wrong answers scatter while the correct answer accumulates agreement.
Jane: Exactly. And that’s why self-consistency, which resamples the same model over and over, inherits that model’s systematic errors. If the model has a blind spot, every sample will have that same blind spot, and the majority vote will confidently pick the wrong answer.
Tom: So the jury is a label-free, training-free verifier. You don’t need annotated data, you don’t need to train a reward model. You just need a few different models that you probably already have access to. That’s a big deal for practitioners.
Jane: It is. And the paper goes further. They derive a mathematical law that predicts when the jury will work and when it won’t. That’s the part we should dig into next, because it gives you a ceiling, a floor, and a way to know in advance whether you can trust the verdict.
Tom: I’m hooked. Let’s get into that law and what it means for actually using this in practice.
Summary: Jane: So Tom, we’ve established the basic idea of the LLM-jury. Now let’s talk about what the paper actually found across their seven benchmarks. They tested on competition math like AIME, on MATH-five hundred on OlympiadBench, on grade-school math, on graduate science with GPQA, and on broad knowledge with MMLU-Pro.
Tom: And the results are pretty consistent. The jury beats self-consistency on every benchmark that has headroom. On AIME-two thousand twenty-four it’s a twenty-point jump, from thirty-six point seven to fifty-six point seven percent accuracy. That’s enormous. And it beats the single-model LLM-verifier on all seven benchmarks.
Jane: The single-model verifier is the weakest selector, which is really telling. That’s the setup where one model scores each candidate solution and picks the highest score. The paper’s point is that a model cannot adjudicate its own candidates, because it shares the same blind spots that produced the errors in the first place.
Tom: Right. And the advantage holds even when they switch the generator. They repeated the whole comparison with DeepSeek-V3 point 2 as the generator instead of Qwen3-235B, and the jury still wins on every unsaturated benchmark. So it’s not about one specific model’s pool of candidates.
Jane: Now, the part I find most exciting is the predictive law. The paper derives a closed-form equation that predicts the jury’s accuracy from just three measured statistics: the mean member accuracy, the pairwise error correlation, and the shared-misconception rate. And it predicts consensus accuracy to a mean absolute error of zero point zero three across all seven benchmarks.
Tom: That’s remarkably precise for a parameter-free law. And it also predicts the method’s ceiling, which they call the shared-error floor. That’s the rate at which the whole panel agrees on the same wrong answer. On math, that floor is near zero. But on GPQA science questions, it’s about three percent, because independently trained models share some graduate-level misconceptions.
Jane: So the law tells you when to trust the jury and where it will fail. That’s the kind of characterization you rarely get for a heuristic method. It’s not just “voting helps.” It’s a precise statement about the conditions under which voting helps and the irreducible error that remains.
Tom: And that irreducible error is a real thing. They found cases on GPQA where all four models were unanimously wrong, and they were wrong for the same reason. For example, all four reported a thermochemical quantity in the textbook-default unit instead of the per-gram basis the question asked for. That’s a shared convention, not a random slip.
Jane: That’s a beautiful example, Tom. It shows the floor is structured, not noise. And it means the jury is blind exactly where models have internalized the same convention. The law quantifies that blindness in advance.
Tom: So we have a verifier that’s free, label-free, and predictable. But the big question for me is how it stacks up against the trained reward models that people are actually using. That’s the comparison that matters for adoption.
Improvements: Tom: So Jane, we’ve seen the jury beat self-consistency and the single-model verifier. But the real test is against trained reward models. The paper compares against four of them: two discriminative process reward models, Qwen2 point 5-Math-PRM at 7B and 72B, an outcome reward model called AceMath-72B, and a generative verifier called ThinkPRM-14B.
Jane: And the results are striking. On MATH-five hundred which is exactly the kind of math these reward models were trained on, the jury matches the strongest of them. It gets ninety-six point six percent, versus ninety-five point six for AceMath. That difference isn’t statistically significant, but the point is the free jury is right there with the best trained verifier.
Tom: And out of domain, on GPQA, the jury is the top selector. It gets thirty-five point nine percent, beating all four trained verifiers, some of them significantly. That’s the key advantage. Trained reward models don’t transfer well outside their training distribution, but cross-model agreement is domain-agnostic.
Jane: That’s the improvement the paper suggests, in a nutshell. Instead of spending resources training a reward model that only works on math, you can use a panel of models you already have. The jury is immediately deployable, with no training cost, and it generalizes to new domains.
Tom: And there’s a clever engineering trick too. They propose an agreement-gated cascade. You run a cheap two-model panel on every problem. If they agree, you accept their answer. If they disagree, you escalate to the full four-model panel. This recovers the full panel’s accuracy at close to two-model cost.
Jane: The escalation rate self-scales with difficulty. On GSM8K, only three percent of problems escalate. On AIME, it’s sixty to seventy percent. So you spend the extra compute exactly where it matters. At matched average cost, this cascade beats self-consistency by a wide margin on AIME-two thousand twenty-four seventy percent versus thirty-six point seven.
Tom: And the signal generalizes beyond math. On code, they define consensus behaviorally. Two programs agree if they produce identical outputs on a shared battery of test inputs. That predicts correctness at an AUROC of zero point seven four eight, and unanimous programs pass held-out tests ninety-five point two percent of the time versus fifty-two point nine percent when the panel splits.
Jane: So the same principle works across modalities. And they even show it can be used as a label-free reward for fine-tuning. A policy trained with cross-model consensus as the reward captures most of the gain that gold labels provide, at both 1 point 5B and 7B scale.
Tom: That’s a full toolkit, Jane. Verification, abstention, routing, and training signal, all from the same simple idea. But I want to hear what Lu and Meng think about the practical side. Lu, does this change how you think about building reasoning systems?
Lu: It absolutely does, Tom. The paper’s framing of a “shared-error floor” is the part I find most profound. It tells us that the ceiling of this method isn’t a property of the method itself, but of how alike current models are. If we train models that fail differently, the floor drops and the jury gets stronger. That’s a research direction in itself.
Meng: And from an engineering standpoint, the cost story is what sells it. You’re not hosting a separate reward model. You’re using the models you already query. The cascade makes the compute story even better. I’d want to see how this holds up with more diverse panels and different model families, but the open-weight results in the appendix are encouraging.
Jane: Great points from both of you. The paper really does span theory, experiments, and practical deployment. Let’s wrap up with our final thoughts.
Conclusion: Tom: We’ve covered a lot of ground on “LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning.” Let’s pull it together. The core idea is that agreement among independently trained models is a powerful verification signal, one that’s free at inference time and needs no training data.
Jane: And the evidence is strong. On competition math, the jury closes the entire gap to an oracle selector. It matches the strongest trained reward models inside their training domain and beats them outside it. And the parameter-free law predicts its behavior to within zero point zero three mean absolute error.
Tom: The law also gives you the method’s ceiling, the shared-error floor. On math, that floor is near zero. On science, it’s measurable, because models share some misconceptions. That’s not a flaw in the method, it’s a measurement of how alike current models are.
Lu: And that measurement points to a concrete path forward. If we want stronger verification, we should train models that fail differently. The law quantifies the value of that investment.
Meng: From my side, the practical takeaway is that you can deploy this today with models you already have. The cascade keeps the cost down, and the abstention dial gives you a calibrated way to decide when to answer and when to escalate.
Jane: So we’ll say goodbye to this paper with a sense of excitement. It’s rare to find a method that’s simple, free, predictable, and competitive with trained systems. The LLM-jury is all of those things.
Tom: And the implications go beyond verification. The same signal can train models, route compute, and even work for code. That’s a versatile tool. Thanks for joining us, and we’ll see you next time with another paper.
Ning Liu
cs.LG, cs.AI
Submitted: 2026-08-17
Updated: 2026-08-18
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 94/100
The gist: The paper introduces the LLM-jury, a training-free verification method for selecting correct answers from a pool of candidate reasoning chains in large language models.
Key concepts
- LLM-Jury
- A verification method where several different large language models solve a problem independently. The final answer is determined by taking a vote or consensus among all the models, suggesting that agreement signals correctness.
- Cross-Model Consensus
- The core idea that the signal separating correct answers from wrong ones exists between multiple models, not within any single model. This 'voting' approach leverages diverse perspectives to catch errors.
- Process Reward Models (PRMs)
- Trained models designed to score or evaluate each step of a solution process. The paper argues that the simple cross-model consensus method often outperforms these complex, trained reward systems.
- Shared-Error Floor
- A mathematically predicted ceiling for the jury's accuracy, representing the rate at which all panel models agree on an incorrect answer due to shared misconceptions or conventions.
Terminology
Summary
The paper introduces the LLM-jury, a training-free verification method for selecting correct answers from a pool of candidate reasoning chains in large language models. The core idea is cross-model consensus: several independently trained models, each solving a problem once, agree on a final answer, and the structure of their agreement serves as the verification signal. The authors state: We treat the panel as an LLM-jury, in which the structure of agreement, not any model's score of another, is the verification signal.
The paper argues that the bottleneck in test-time scaling has shifted from generation to selection: "As the sample count grows a correct answer almost always appears somewhere in the pool, so the bottleneck shifts from generation to selection: the verifier that picks which candidate to return, not the generator, increasingly bounds end-to-end accuracy." The two dominant verifiers each carry structural costs: self-consistency inherits the errors of the single model it resamples,
and trained reward models need labeled data and transfer poorly off-distribution.
The LLM-jury operates as follows: a panel of M independently trained models each produce a final answer to a problem... The cross-model consensus answer is the modal class's value, and the agreement g ∈ (0, 1] is the fraction of models in that class.
Crucially, our jurors never see the candidates or one another's work,
unlike LLM judges that read and score generations.
The authors compare four selectors: self-consistency (one model sampled N times, majority vote), cross-model consensus (panel majority over M models), single-model LLM-verifier (one model scores each candidate), and oracle (selects any correct candidate). All selectors operate on the same fixed candidate pool so differences reflect the scoring signal, not the candidates.
1. Cross-model consensus is a strong Best-of-N verifier. On a fixed pool of N=12 candidates from Qwen3-235B, the LLM-jury is the strongest non-oracle selector on every unsaturated benchmark.
It beats self-consistency significantly on every benchmark with headroom (up to +20.0 on AIME-2024, p=0.001) and beats the single-model LLM-verifier on all seven benchmarks. The single-model verifier captures essentially none of the oracle gap above self-consistency,
providing direct evidence that a model cannot adjudicate its own outputs.
2. A parameter-free predictive law. The authors derive a closed-form law predicting consensus accuracy and its ceiling from three measured panel statistics: mean member accuracy a, mean pairwise error correlation ρ, and shared-misconception rate s. The law matches empirical consensus accuracy to a mean absolute error of 0.028 and the shared-error floor to 0.009
across seven benchmarks with no per-benchmark fitting. The shared-error floor—the rate at which the whole panel agrees on the same wrong answer—is near zero on competition math, where wrong answers scatter, but 0.030 on GPQA, where independently trained models share graduate-science misconceptions.
3. The jury outperforms trained verifiers. Against four trained verifiers (Qwen2.5-Math-PRM 7B/72B discriminative PRMs, AceMath-72B-RM outcome RM, ThinkPRM-14B generative verifier), the LLM-jury is the strongest non-oracle selector on all four benchmarks, at no training cost.
On MATH-500, the jury (0.966) "significantly exceeds the two discriminative PRMs and the generative verifier (0.928–0.946; paired bootstrap p≤0.014) and marginally exceeds the strongest trained verifier, the outcome RM AceMath-72B (0.966 vs. 0.956, not significant at p=0.142)." Out of domain on GPQA, it is again the top selector (0.359 vs. 0.268–0.323).
4. Error decorrelation is the mechanism. The authors measure pairwise error correlation: within-model (self-consistency, ρ̄=0.68) > same-family (0.52) > cross-family (0.47). They show a panel of distinct models beats one model resampled as often (+12.5 to +16.7 on AIME-2024), so the gain is decorrelation rather than sample count or raw strength.
Furthermore, single-model accuracy plateaus by N≈8
while a four-model panel reaches 70.0% on AIME-2024 with 4 generations, an 8× smaller budget yet +26.7 points.
-
Agreement as abstention dial:
Answering only when the panel is unanimous sharply raises accuracy: on MATH-500 a unanimous four-model panel is correct 99.5% of the time while still answering 85% of problems.
-
Agreement-gated cascade: Running a cheap two-model panel and escalating to the full four-model panel only on disagreement
recovers the full-panel accuracy at close to two-model cost.
-
Code domain: On HumanEval+, behavioral consensus (programs agreeing on shared test inputs)
predicts correctness at AUROC 0.748,
with unanimous programs passing held-out tests 95.2% of the time versus 52.9% when the panel splits. -
Consensus as training reward: As a label-free RFT reward,
cross-model consensus recovers most of the gain that gold labels provide
(at 1.5B: +4.5 vs. +5.5 gold; at 7B: +8.5 vs. +10.5 gold). -
Robustness: The advantage holds with a different generator (DeepSeek-V3.2), with a fully open-weight panel, and with a frontier-model panel.
The authors note: "The LLM-jury inherits a hard ceiling at the shared-error floor: where independently trained models share a misconception (measurably, parts of GPQA), unanimous agreement is confidently wrong and no agreement-based verifier can help; selection likewise cannot exceed the pool's oracle. As a voting rule it does not beat the single strongest panel member when panel strength is highly unequal... It also requires more than one model, which a single-model deployment cannot supply."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:
-
Add a cross-model consensus verifier module. Instead of relying on self-consistency (resampling one model) or a trained reward model, the system will, at inference time, call 3–4 independently trained models (e.g., Qwen, DeepSeek, Claude, Kimi) to solve the same problem once each. The system selects the modal answer class as the final output.
-
Implement the parameter-free predictive law for abstention. The system will measure three panel statistics per domain (mean member accuracy
a, pairwise error correlationρ, shared-misconception rates) on a small labeled calibration set (≈200–500 problems). Using the closed-form law (Eq. 1 and Eq. 2), it will predict, in advance, the consensus accuracy, the full selective-prediction curve, and the shared-error floor. This lets the system decide before deployment whether to trust the jury on a given domain. -
Add an agreement-gated cascade for compute efficiency. The system will run a cheap two-model panel on every problem. If the two agree, accept their answer (cost: 2 calls). If they disagree, escalate to the full four-model panel (cost: 4 calls). This recovers full-panel accuracy at near-cheap cost, with escalation rate self-scaling with difficulty (3% on GSM8K, 60–70% on AIME/GPQA).
-
Use agreement fraction as a calibrated abstention dial. The system will return the answer only when the agreement level
gexceeds a threshold. For example, on MATH-500, a unanimous four-model panel is 99.5% correct at 85% coverage; on AIME, unanimous is 100% correct. The system will set the threshold based on the predicted selective-prediction curve. -
Replace trained reward models with the jury for out-of-domain tasks. The system will use the jury as the default Best-of-N selector, especially where no in-domain reward data exists (e.g., GPQA, code, heterogeneous tasks). The jury matches or exceeds trained PRMs (Qwen2.5-Math-PRM 7B/72B, AceMath-72B, ThinkPRM-14B) inside their math domain and is the top selector outside it, at zero training cost.
-
Use behavioral consensus for code generation. For code tasks (e.g., HumanEval+), the system will define agreement behaviorally: two programs agree if they produce identical outputs on a shared battery of test inputs. Programs in a unanimous behavioral class pass held-out tests 95.2% of the time versus 52.9% when the panel splits.
-
Use cross-model consensus as a label-free training reward. For rejection-sampling fine-tuning, the system will keep solutions that match the cross-model peer consensus (not self-consistency, not gold labels). On a 7B policy, this recovers ≈81% of the gold-label gain on MATH-500 (+8.5 vs. +10.5 points), with no labels.
-
Select correct answers from a candidate pool better than self-consistency and single-model verifiers. On AIME-2024, it captures 100% of the oracle gap (56.7% vs. 36.7% for self-consistency, p=0.001). On OlympiadBench, +7.2 points; on GPQA, +2.6 points.
-
Match or beat trained reward models at zero training cost. On MATH-500, it ties the strongest trained verifier (AceMath-72B, 96.6% vs. 95.6%, p=0.142) and beats three others significantly (p≤0.014). On GPQA, it is the top selector (35.9% vs. 26.8–32.3%).
-
Predict its own reliability in advance. The law predicts consensus accuracy to a mean absolute error of 0.028 and the shared-error floor to 0.009, with no per-benchmark fitting. The system can thus decide before deployment whether to trust the jury on a new domain.
-
Abstain or escalate based on calibrated agreement. It can answer only when the panel is unanimous (99.5% correct on MATH-500 at 85% coverage; 100% on AIME) or escalate to more compute only on disagreement, matching full-panel accuracy at near-cheap cost (2.2 calls/problem on MATH-500 vs. 4).
-
Work fully open-weight. A three-model open panel (Qwen3-235B, DeepSeek-V3.2, Kimi-K2.5) reproduces the full advantage, so the system does not depend on proprietary APIs.
-
Improve itself without labels. Using consensus as a reward, the system can fine-tune a 7B policy to gain +8.5 points on MATH-500 (transfer from GSM8K), nearly matching the +10.5 from gold labels, with no ground-truth annotations.
-
Detect and bound its own failure mode. The system will know its shared-error floor (e.g., ≈0.03 on GPQA, ≈0.14 on MMLU-Pro) and can flag domains where unanimous agreement is likely wrong (shared misconceptions in chemistry/physics), rather than silently returning a confidently wrong answer.
Abstract
Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution. We study a third signal, free at inference time: cross-model consensus, the degree to which independently trained models, each solving the problem once, agree on a final answer. We treat the panel as an LLM-jury, in which the structure of agreement, not any model's score of another, is the verification signal. Across seven benchmarks it selects correct answers better than self-consistency and far better than a model scoring its own candidates: on competition math it closes the entire gap to an oracle selector, while self-scoring closes almost none. The mechanism is error decorrelation: independently trained models err differently, so their wrong answers scatter while the correct one accumulates agreement. We make this precise with a parameter-free law, derived in closed form, that predicts consensus accuracy from three measured panel statistics to a mean absolute error of 0.03 and exposes the method's ceiling: a shared-error floor where models share a misconception, near zero on math but non-trivial on science. Against four trained verifiers spanning discriminative, outcome, and generative reward models, the free LLM-jury matches the strongest inside their math training domain and is the top selector outside it. Cross-model consensus is thus a verifier we can characterize in advance: a law that says when to trust it, and a floor that marks where it cannot.
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Universal Self-Consistency for Large Language Model Generation
- Training Verifiers to Solve Math Word Problems
- Large Language Models Cannot Self-Correct Reasoning Yet
- Language Models (Mostly) Know What They Know
- Process Reward Models That Think
- Let's Verify Step by Step
- Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
- AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
- s1: Simple test-time scaling
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Mixture-of-Agents Enhances Large Language Model Capabilities
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
- The Lessons of Developing Process Reward Models in Mathematical Reasoning
- Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks