JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation

arXiv:2609.40103 · cs.CL, cs.HC · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation".

Tom: Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet when judges disagree, majority voting discards this conflict instead of resolving it.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So to wrap up what we've heard about "JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation," the authors are proposing a system that uses disagreement as a signal to guide human intervention rather than just letting majority voting discard those conflicts.

Jane: It’s essentially moving away from simple aggregation when judges don't agree, choosing instead to pinpoint the exact source of uncertainty.

Meng: The title itself suggests this is a framework built around handling that messy disagreement in an automated way, focusing on targeted human input.

Lu: The implication here is that we can make our multi-agent evaluation systems much more robust by systematically resolving where the underlying judgment standards are shaky.

Tom: I think what this means in plain terms is that instead of just getting a consensus score when judges argue, we get a clear map of what's causing the argument and fix that specific piece so everyone learns correctly moving forward.

Jane: That moves the evaluation from being a static scoring exercise to an active learning process guided by human insight into complexity.

Lalam: For us at this startup, this is huge because it means the AI isn't just producing content; it's helping us build the reliable knowledge base that dictates what good content looks like.

Tom: The paper explores how to move from a brittle system that fails when judges disagree to one that uses those conflicts as precise indicators of where we need help.

Jane: It shows the value in decomposing responses into claims so we can focus our limited human time on the most critical points of contention.

Lu: This approach shows a path toward creating evaluation systems that are more aligned with nuanced human preferences across different judges and domains.

Tom: We're looking at how this structure, using the disagreement graph and propagated re-evaluation, can create a feedback loop that makes all agents better with every single piece of content they judge.

Meng: It’s about building a mechanism for continuous self-improvement within the evaluation pipeline itself.

Lalam: Ultimately, JuryFlow suggests that by treating disagreement as information, we can create much more sophisticated and reliable AI content assessment tools than what we have now.

Conclusion: Tom: So, we've been deep in JuryFlow, and now we need to talk about what that title actually means for us as listeners. Jane It’s all about how this new framework uses disagreement as a guiding force instead of just ignoring it. Lu That title really captures the core idea of treating conflict between judges not as noise, but as a signal pointing toward where the AI evaluation is uncertain.

Meng: From an engineering standpoint, that shift from averaging out to pinpointing uncertainty sounds like a significant step in making our evaluation pipeline more transparent and traceable. Tom Exactly! It moves us from getting a single, potentially misleading consensus score to understanding exactly *why* the judges couldn't agree on certain claims. Jane And the authors are doing this by creating a system where human input is strategically targeted only to resolve those specific points of disagreement.

Lalam: For me, seeing this focus on structural similarity and propagating corrections across instances suggests a way to build an evaluation culture that learns from its own mistakes very quickly. Lu I think the real power here lies in how it crystallizes those corrected judgments into reusable rubrics that all our agents can inherit later. Tom That’s a big deal because it means the learning isn't just temporary; it becomes baked into the system's DNA for future content checks.

Jane: It feels like we are building a system that doesn't just pass tests, but actively improves its own judgment process based on where it falters. Meng And when you think about the broader impact, this could lead to AI systems that are far more reliable and less prone to making subtle errors in complex tasks.

Tom: It really puts the focus squarely on improving the *process* of judging rather than just tweaking the final output score. Lu We've seen how this architecture handles adversarial cases better because it knows exactly where those tricky disagreements are originating. Jane So, JuryFlow seems to be about building smarter judges, not just smarter content generators.

Lalam: Precisely. If we can consistently feed these refined rubrics back into our models, the overall quality and trustworthiness of the AI-generated material will see a real upward trend across the board. **Transition Sound**

Mufeng Yang, Junwei Yu, Yepeng Ding

University of Tsukuba · The University of Tokyo · Hiroshima University

cs.CL, cs.HC

Submitted: 2026-09-30

Updated: 2026-09-30

Importance score: 91/100

The gist: Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet when judges disagree, majority voting discards this conflict instead of resolving it.

Key concepts

Atomic Claims
This process breaks down an entire AI response into its smallest, individual statements or claims. Each claim is then given a verdict by different judges and assigned relevant tags. This decomposition allows the system to pinpoint exactly where disagreements occur among the judges.
Disagreement Graph
A map built from the atomic claims where each claim is a node. The connections (edges) between nodes show how similar or related two claims are, based on shared tags or semantic similarity. This graph helps track how uncertainty flows from one claim to another.
Human-in-the-Loop Intervention
Instead of having a human re-evaluate the entire response, the system presents the disagreement graph and asks for a single choice: which specific claim to resolve. This minimal intervention focuses human effort where it is most needed, guiding the subsequent automated correction process.
Reusable Rubric Entry
After correcting a focal claim, JuryFlow generates a standardized entry detailing the correct evaluation criterion. This entry includes natural language rules and examples, which are then inherited by all judge agents in future evaluations to ensure consistent scoring.

Terminology

Summary

Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet when judges disagree, majority voting discards this conflict instead of resolving it. This paper presents JuryFlow, a disagreement-guided, human-in-the-loop multi-agent evaluation framework that treats inter-judge disagreement not as noise to be averaged away, but as a precise signal indicating where an evaluation is uncertain.

The gist

JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph whose nodes are scored by verdict entropy and whose edges encode structural similarity between claims. A human acts as a structural guide, i.e., selecting which disagreement to resolve through a single, minimal intervention rather than re-labeling the response, after which the focal claim is re-evaluated, the correction propagates along graph edges and to historically similar cases, and is crystallized into a reusable rubric that all agents inherit in subsequent evaluations.

How it works

JuryFlow operates in five stages designed to treat disagreement as an evaluation signal. In Stage 1, a panel of heterogeneous judge agents independently decomposes a candidate response into atomic claims and assigns per-claim verdicts and dimension tags. This stage involves obtaining model diversity primarily by using LLMs from different providers, and optionally adding prompt diversity by varying evaluation personas or domain-specific instructions.

In Stage 2, the system constructs a disagreement graph where each node corresponds to a claim. Node-level disagreement score is computed based on the entropy of the verdict distribution across agents: d i = -∑ v p i(v) log p i(v). Claims with scores exceeding a threshold τ are marked as disagreement candidates. Edge weights capture structural similarity between claims, defined by: w ij = α · simtag (c i, c j) + (1 − α) · simemb (c i, c j), where simtag is the Jaccard similarity over the union of agent-assigned tags and simemb is the cosine similarity of claim embeddings.

How it works

Stage 3 involves human intervention. The system presents the disagreement graph and highlights top-m candidates ranked by entropy (or selects them automatically via entropy ranking in experiments). The human makes one selection instead of re-evaluating the full response, choosing a focal claim c∗ to resolve. For automatic benchmarking, this selection is made by entropy ranking.

Stage 4 triggers propagated re-evaluation based on the selected focal claim c∗. This involves two processes: Focal re-judgment, where a reviewer agent re-evaluates the focal claim c∗ using supporting and opposing evidence, and Graph-based propagation. Intra-instance propagation flags claims connected by edges with weight w > β for automatic re-evaluation, while Cross-instance propagation uses nearest neighbor embedding search to retrieve historical instances with similar high-disagreement claims. The threshold β controls the trade-off between correction coverage and computational cost.

How it works

Stage 5 converts the focal correction into a reusable rubric entry. This involves analyzing c∗, its tags, original verdicts, and the revised verdict to identify the evaluation dimension and boundary condition associated with the disagreement. An LLM then produces a candidate entry containing a natural-language criterion, one positive example from the corrected case, and one negative example from the original erroneous judgment. All judge agents inherit this updated rubric in subsequent evaluations.

Experimental Setup

JuryFlow is evaluated on MT-Bench and LLMBar using an automatic selection configuration where the focal claim is chosen by entropy ranking. The panel comprises N=5 heterogeneous judge agents: GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3.1-70B-Instruct, and Qwen2.5-72B-Instruct to ensure model diversity through distinct LLM families and prompt diversity.

The evaluation protocol uses pairwise comparisons for MT-Bench and LLMBar instances. The final response quality score is reduced to a response-level score using a weakest-link rule, which penalizes the response in proportion to its rejected and uncertain claims. Baselines include a single judge, majority-vote panel, and JuryFlow (full), comparing agreement with gold labels and Cohen’s κ. Cost is quantified by reporting the mean re-evaluation LLM calls per instance.

Key Findings

JuryFlow improves over the best single judge by 5.3 accuracy points on MT-Bench and by 10.5 points on LLMBar, and exceeds the stronger majority-vote panel by 2.9 and 6.6 points, respectively. The gain is larger on LLMBar, which contains adversarial cases prone to judge disagreement. Ablations show that focal re-evaluation alone lifts the majority-vote baseline by 3.

Improvements for AI systems

Based on the JuryFlow framework, here are specific improvements that can be made to existing AI evaluation systems, and what those improved systems could achieve:


  1. A system capable of performing disagreement-guided quality assurance on complex AI outputs by treating inter-judge conflict not as noise to be averaged out, but as a precise signal for targeted refinement.

  2. The ability to decompose large AI responses into atomic claims and assign per-claim verdicts from a heterogeneous panel of LLMs (e.g., GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro).

  3. The capability to construct a Disagreement Graph where nodes are claims scored by verdict entropy (measuring the strength of conflict) and edges encode structural similarity between claims (based on both tag Jaccard similarity and claim embeddings).

  4. A mechanism to identify the single most critical disagreement point—either through human selection or automatic selection via entropy ranking—to resolve. This replaces expensive, full-response re-evaluation with a highly focused intervention.

  5. The ability to perform Propagated Re-Evaluation: once a focal claim is corrected, this correction automatically propagates along the graph to structurally related claims and cross-instance retrieval of historically similar cases, ensuring consistency across the entire corpus or task set.

  6. The capacity for Incremental Rubric Induction: every targeted correction is abstracted into a reusable rubric entry (a natural language criterion with positive/negative examples), which is then inherited by all future agents in subsequent evaluations, leading to a progressively self-refining evaluation system.

This improved AI system can achieve the following specific capabilities:

  1. It can significantly increase the agreement accuracy between different AI judges compared to single-judge or majority-vote panels (as shown on MT-Bench and LLMBar).

  2. It provides a more robust mechanism for identifying high-quality, actionable errors in complex AI outputs by focusing computational resources only where consensus is weakest.

  3. It drastically reduces the computational cost of quality assurance by replacing full response re-evaluation with targeted, evidence-grounded re-judgments, while still achieving higher accuracy (as evidenced by lower Calls metrics in ablation studies).

  4. It creates a knowledge base of refined evaluation criteria that evolves over time, allowing the system to learn and adapt to specific failure modes across different instances without requiring complete retraining of the judging panel.

  5. It can automatically detect and flag subtle errors that might be missed by simple confidence-based aggregation, by prioritizing claims where the panel is most evenly divided (highest entropy).

Sources

Related papers