Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

arXiv:2608.12585 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar

Amazon AGI

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces Authors: Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar (Amazon AGI) Summary This paper

Terminology

Summary

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

Summary

This paper introduces Reasoning Jury, a system designed to improve the fidelity of judgments for identifying defects in long reasoning traces produced by reasoning LLMs. The core idea is to replace a single LLM judge with a jury of multiple LLMs and a moderated consensus mechanism. The system is motivated by the observation that single-model judges, even frontier models, struggle to identify reasoning defects in long, complex chain-of-thought traces. Additionally, using frontier models for online training is often prohibited by guardrails, making a high-fidelity jury of open-weight models a valuable alternative.

Problem and Motivation

The paper argues that reliable evaluation of reasoning traces is critical for several stages of LLM development:

  • Supervised fine-tuning: Filtering low-quality reasoning data.

  • Reinforcement learning: Providing reward signals that distinguish good from bad reasoning.

  • Performance evaluation: Diagnosing reasoning outputs and understanding failure modes.

Unlike short-answer evaluation, judging long reasoning traces requires identifying which reasoning step went wrong and how badly. A single bad step can propagate through the entire trace. The paper states that single-model judges (even frontier models) do not do well at identifying reasoning defects on this complex task.

The Reasoning Jury System

The system operates on reasoning traces that are segmented into steps, demarcated by [STEP-x] markers. The segmentation is performed using a regex pattern that identifies an end-of-sentence character followed by a double newline ((?<=[.!?]) +), which is more robust than simple double-newline splitting.

The pipeline consists of two phases:

  1. Phase 1: Independent Judgement: Each juror LLM independently reads the problem, the step-annotated trace, and the final solution. It emits a structured list of defects. Each defect includes:
  • what went wrong: A self-contained description of the defect.

  • statement refs: A list of [STEP-x] markers where the defect occurs.

  • impact: Severity (neutral, minor, major, fatal).

  • evidence: Direct quotes from the trace supporting the defect.

  • The prompt enforces a genericness test and a specificity self-check to suppress vague criticism.

  1. Phase 2: Consensus: A consensus is derived from the Phase 1 verdicts using one of two modes:
  • Consolidation: A judge (moderator) is given the original content and all Phase 1 verdicts and is asked to verify, merge, and fill gaps in the judgements. The final output is in the same format as Phase 1 outputs.

  • Deliberation: A moderator runs a multi-turn debate among the jurors. The moderator is intentionally blind to the problem, reasoning trace, and solution, seeing only the deliberation transcript. Its role is purely procedural: managing turns, surfacing disagreements, and detecting convergence. The loop terminates on a full round with no new arguments, universal agreement, or detected cycling. A final extraction call synthesizes the consensus defects, a confidence score, and dissenting views.

Key Results and Findings

The system was evaluated on the Hard2Verify and DeltaBench benchmarks, focusing on step-level defect localization using Balanced F1 score.

  • Single-Model Baseline: On Hard2Verify, gpt-5.4 was a saturated outlier at 83.9 Balanced F1, far ahead of other single models like opus-4.6 (73.7), sonnet-4.6 (71.2), and gemini-3.1-pro (69.5).

  • Jury vs. Single Judge: A jury of open-weight models significantly outperformed frontier models.

  • The Small/Medium OSS jury (including gpt-oss-120b, qwen3.6-27b, etc.) reached 80.8 Balanced F1 with deliberation, outperforming opus-4.6, sonnet-4.6, and gemini-3.1-pro by 7.1–11.3 points.

  • The Large OSS jury reached 80.2 Balanced F1.

  • The full Frontier jury (including gpt-5.4) reached 84.4, but this was near-saturated by gpt-5.4's solo score, showing that a jury adds little when one member is dominant.

  • Deliberation generally yielded higher Balanced F1 than consolidation, primarily by improving recall.

  • Jury Size: Experiments showed that three jurors provide a reasonable operating point for most panels, with performance declining more sharply when reduced to two.

  • Moderator Choice: The choice of moderator matters. With deliberation, moderator choice moved consensus by 2.6–3.7 points; with consolidation, the range was larger (6.1–6.4 points), showing consolidation is more sensitive to the moderator's capability.

  • Diversity Control: Homogeneous juries (three independent samples of the same model) also showed large gains. A 3× gpt-oss-120b jury reached 82.3 Balanced F1, a +13.0 lift over its best solo score. This shows that model diversity is not necessary for gains; independent sampling and deliberation are key, though the ceiling is still constrained by base-model capability.

  • Generalization to DeltaBench: A homogeneous 3× gpt-oss-120b jury with deliberation reached 61.5 Balanced F1, outperforming a single opus-4.6 (58.4) and matching gpt-5.4 (61.8).

  • Cost Analysis: The jury is significantly cheaper than using a frontier model as a judge.

  • A single opus-4.6 pass on 200 Hard2Verify records cost 79.13.

  • The 3× gpt-oss-120b jury cost ** 12.35** in deliberation mode (6.4× cheaper) and ** 6.50** in consolidation mode (12.2× cheaper).

  • This is despite using 7.6× more tokens and 10.9× more LLM calls, because the per-token price of gpt-oss-120b is much lower than opus-4.6.

  • Comparison to Naive Aggregation: Majority voting performed poorly due to sparse exact-step agreement. The union of defects captured most of the accuracy benefit, but deliberation and consolidation provide a deduplicated, adjudicated, and structured verdict, which is more useful for downstream applications.

Exploratory Output Profiling

The paper demonstrates the jury's use as an analytical instrument by profiling nemotron-3-super on AIME2026 problems. The jury's structured outputs were aggregated into a defect taxonomy, revealing failure modes. Key findings:

  • The most common defect was extended reasoning on an uncorrected false premise, which was also predominantly fatal.

  • The defect signal tracked final-answer correctness: traces with wrong answers had 4.86 defects per trace (2.31 fatal), while correct-answer traces had 0.99 defects per trace (0.17 fatal). This exposes right answer, flawed reasoning cases.

Downstream Use: Defect-Guided Retry

A task-based evaluation showed that the jury's feedback helps a model correct its reasoning. For 877 flagged traces from nemotron-3-super:

  • Providing only step locations led to 71.2% retry accuracy.

  • Providing the full rich feedback (what went wrong diagnosis and severity) raised retry accuracy to 76.2%.

  • This rich feedback also caused fewer regressions (correct→wrong) than step anchors alone (18 vs. 27).

Limitations

The paper acknowledges several limitations:

  • The jury is slower and consumes more tokens than a single judge, making it less suitable for online use cases like on-policy RL.

  • The benchmark evaluations directly measure defect localization but do not validate the factual accuracy of what went wrong descriptions, severity calibration, or downstream usefulness of these fields.

  • The defect taxonomy used for profiling is a single induction run and is not a canonical or validated standard.

Conclusion

The paper concludes that Reasoning Jury is a practical approach for offline reasoning evaluation when a single judge is insufficient. Independent sampling provides the main accuracy benefit, while consolidation and deliberation turn pooled findings into a usable, structured verdict with different cost-latency tradeoffs. The system enables a jury of cheap, open-weight models to outperform expensive frontier models at a fraction of the cost, providing a viable path for high-fidelity reasoning evaluation and training signal generation.

Improvements for AI systems

Improvements to AI Systems:

  1. Multi-Model Consensus Judging Module
  • Integrate a jury-based evaluation layer into AI training pipelines (SFT/RL) that uses 3+ open-weight models with moderated deliberation to score reasoning traces.

  • The improved system can generate high-fidelity, step-level defect annotations (with severity and evidence) at 6–12× lower cost than frontier single-model judges, enabling scalable reward signal generation for RL without relying on restricted frontier APIs.

  1. Deliberation-Driven Self-Correction Loop
  • Add a blind-moderator debate mechanism where multiple AI instances critique each other’s reasoning, converging on a consensus verdict without seeing the original problem.

  • The improved system can autonomously identify and repair flawed reasoning chains in long outputs, boosting task-level retry accuracy from 71.2% (step-only feedback) to 76.2% (rich feedback), while reducing correct→wrong regressions by 33%.

  1. Defect-Aware Data Filtering for SFT
  • Use the jury’s structured output (defect type, step refs, severity) to filter low-quality reasoning traces before fine-tuning.

  • The improved system can automatically remove traces with fatal false-premise errors, preventing the model from learning harmful reasoning patterns and improving downstream reasoning robustness.

  1. Cost-Effective Offline Evaluation Harness
  • Replace single frontier-model evaluators with a homogeneous jury (e.g., 3× gpt-oss-120b) using deliberation, which matches or beats frontier models (e.g., 82.3 vs. 83.9 Balanced F1 on Hard2Verify) at 6.4× lower cost.

  • The improved system can run large-scale offline evaluation of reasoning models (e.g., on 200+ records) for under 13, enabling frequent regression testing during development.

  1. Diagnostic Profiling for Failure-Mode Analysis
  • Aggregate jury outputs into a defect taxonomy to profile model weaknesses (e.g., “extended reasoning on uncorrected false premise” as top fatal error).

  • The improved system can automatically generate a failure-mode report for any reasoning model, revealing “right answer, flawed reasoning” cases (0.99 defects on correct vs. 4.86 on wrong traces) and guiding targeted fine-tuning or prompt engineering.

  1. Adaptive Consensus Strategy Selector
  • Dynamically choose between consolidation (faster, cheaper) and deliberation (higher recall) based on task criticality and latency budget.

  • The improved system can optimize evaluation pipelines for real-time vs. batch use cases, trading off accuracy (deliberation) for speed (consolidation) without manual tuning.

  1. Diversity-Free Performance Booster
  • Leverage independent sampling of the same model (homogeneous jury) to boost judge performance by +13.0 Balanced F1 over solo scoring, without needing diverse model families.

  • The improved system can upgrade any existing single-judge setup to a jury with minimal integration effort, improving evaluation fidelity for any base model.

  1. Structured Feedback Generator for Downstream Agents
  • Output machine-readable defect lists (with step refs, severity, evidence) that can be directly consumed by agentic systems for iterative refinement.

  • The improved system can provide actionable, localized feedback to AI agents, enabling them to correct specific reasoning steps rather than regenerating entire responses, improving sample efficiency in multi-turn problem-solving.

Abstract

Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation. Additionally, surfacing reasoning mistakes that the model makes would enable improving the model's performance at runtime through providing feedback. Due to the difficulty of this complex task on long reasoning traces, single-model judges (even frontier models) do not do well at identifying reasoning defects. Additionally, leveraging frontier models during online training of reasoning LLMs is generally prohibited due to guardrails in terms of use. In this work, we introduce Reasoning Jury, a system that replaces the single judge with a jury of LLMs and a moderated consensus mechanism, to improve the fidelity of judgments for identifying reasoning defects. In reasoning jury, defects of a reasoning trace and their severity are surfaced through a deliberation where a moderator conducts a discussion amongst the jury where the jurors critique each other's judgments and get to modify their initial votes. The moderator derives a consensus through deliberation amongst jurors or consolidation of judgements. We show that Reasoning Jury with a jury of open-weight models (e.g., gpt-oss-120b) is able to significantly outperform frontier models (opus-4.6, sonnet-4.6, and gemini-3.1-pro) at correctly identifying reasoning defects. Besides accuracy performance improvements, the aggregated cost of the jury (initial verdicts, deliberations, consolidation, etc.) is a fraction (8 to 15%) of the cost of running frontier models in LLM-as-a-judge setup. We also show how these judgements can be leveraged to understand failure modes of reasoning LLMs on benchmarks, which allows much deeper understanding of a model's performance.

Sources

Related papers