Atomic and Holistic LLM Judges for Reference-Grounded Support Labels: A Prompt-Controlled Comparison

arXiv:2603.28005 · cs.CL · Submitted 2026-03-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Atomic and Holistic LLM Judges for Reference-Grounded Support Labels".

Tom: This paper investigates whether breaking down an answer into atomic claims for verification provides a performance advantage over using a holistic rubric when classifying reference support in benchmark-style reference-grounded classification…

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So Jane, we’re looking at this paper titled "Atomic and Holistic LLM Judges for Reference-Grounded Support Labels: A Prompt-Controlled Comparison." It sounds like they’re really digging into how we should structure these judges when we want to classify reference support.

Jane: Exactly, Tom; the title hints that they are directly comparing two different ways of asking the AI to make a judgment: one that breaks things down into small pieces, and another that looks at the whole thing at once.

Lu: I think what’s interesting here is how they handle the core idea of breaking an answer into claims versus just scoring it as a whole; it touches on how our models actually process structured information.

Meng: From an engineering standpoint, I’m curious if this comparison is really about the prompting trick, or if the underlying architecture of decomposition itself gives us more control over accuracy.

Lalam: I think these judges are pretty foundational because they decide what constitutes a successful classification on benchmark tasks like ASQA and QAMPARI.

The paper's summary: Tom: Okay, so the paper summarizes that they tested both the self-decomposing atomic judge and the prompt-controlled holistic judge across three different QA datasets to see which one performs better when classifying reference support as fully supported, partially supported, or unsupported.

Jane: That’s right; they found that on some datasets, specifically ASQA and QAMPARI where completeness is really important, the holistic approach actually matches or even exceeds the performance of the atomic judge.

Lu: The summary highlights a key finding: for these completeness-heavy settings, the holistic judge shows statistically reliable gains in three out of four model families they tested.

Meng: So they're saying that decomposition doesn't automatically give you better accuracy, especially when you’re dealing with completeness standards. That’s a practical point for us to consider when building our evaluation pipelines.

Lalam: It means we can lean on the holistic rubric if our main goal is reliably spotting incompleteness in reference support, which is where they saw the biggest difference.

The paper's improvements: Tom: Now, they point out that their main improvement isn't necessarily a new method, but rather a controlled empirical counterexample to the idea that decomposition always helps. They used things like source-level paired tests and cross-family replication to make sure the results weren't just a fluke from one specific prompt wording.

Jane: It’s interesting because they are showing that matching the rubric properly is what matters, not just forcing the AI into a step-by-step decomposition process. They controlled for things like input matching and rubric similarity to isolate that effect.

Lu: The paper suggests that by focusing on a "matched holistic rubric," we can achieve results competitive with the atomic approach on two of the three benchmarks they tested, and it was even cheaper across all three datasets.

Meng: That token accounting detail is pretty telling; if decomposition meant more tokens without a proportional accuracy gain, that’s a real efficiency win for our systems.

Lalam: So the core suggestion is that for completeness tasks, we should be looking at the holistic rubric as a strong alternative to forcing every single claim to be broken down before we judge.

Conclusion: Tom: To wrap things up, the main takeaway from "Atomic and Holistic LLM Judges for Reference-Grounded Support Labels: A Prompt-Controlled Comparison" is that for completeness benchmarks like ASQA and QAMPARI, a matched holistic rubric is competitive with the self-decomposing atomic judge.

Jane: That’s right; the paper shows that while atomic judges have a small edge on TruthfulQA, the holistic design wins on detecting incompleteness in those more demanding settings.

Lu: It suggests that we should probably favor the holistic approach when our primary concern is reliably identifying when an answer is only partially supported by its reference.

Meng: We need to keep in mind their caveat: they found that reference-quality degradation caused the largest drops in accuracy for both designs, so we can't just rely on these judges if our reference data is messy.

Lalam: So, essentially, we’re getting a solid comparison here showing that matching the holistic rubric can be as good as decomposing everything when you're focused on completeness.

Xinran Zhang

University of California, Berkeley

cs.CL

Submitted: 2026-03-30

Updated: 2026-09-30

Importance score: 92/100

The gist: This paper investigates whether breaking down an answer into atomic claims for verification provides a performance advantage over using a holistic rubric when classifying reference support in

Key concepts

Atomic Judge
This method instructs an LLM to first break a candidate answer into distinct factual claims. The judge then verifies each claim individually, determining if it is supported or unsupported. This approach aims to be precise by checking every piece of information separately.
Holistic Judge
This design asks the LLM to score the entire candidate answer using a comprehensive rubric covering various aspects like correctness and completeness, without forcing an explicit claim-by-claim breakdown. It relies on the model's ability to assess the overall quality of the support.
Completeness-Sensitive Benchmarks
These are classification tasks where accurately identifying missing information is crucial for a correct label. Datasets like ASQA and QAMPARI are examples, making them challenging because simply being mostly correct isn't enough; you must detect what is absent.
Token Accounting
This involves measuring the length of the generated text (tokens) produced by each judge design. The study found that atomic judges generate substantially more tokens than holistic judges on certain benchmarks, suggesting decomposition doesn't provide proportional accuracy benefits.

Terminology

Summary

This paper investigates whether breaking down an answer into atomic claims for verification provides a performance advantage over using a holistic rubric when classifying reference support in benchmark-style reference-grounded classification tasks. It compares two judge designs—a self-decomposing atomic judge against a prompt-controlled holistic judge—across three distinct QA datasets, finding that the holistic approach is competitive or superior on completeness-sensitive benchmarks like ASQA and QAMPARI, particularly in detecting incompleteness.

Judge Designs Compared

The study compares two primary judge designs:

  1. Atomic Judge: This design instructs the model to break the candidate into atomic factual claims, identify supported and unsupported claims, and return a structured verdict. This is implemented via various prompt variants (v1, v2, v3), with v3 emphasizing a strict step-by-step protocol for claim verification.

  2. Prompt-controlled Holistic Judge: This design asks the model to score the candidate using a detailed rubric covering correctness, completeness, unsupported detail, and resistance to style bias without requiring explicit decomposition into claims.

Experimental Setup and Comparison

The comparison was rigorously controlled across several dimensions to isolate the effect of the judge design rather than prompt wording or model capability. The methodology included:

(1) Prompt Control:

(2) Input Matching:

(3) Detailed Rubric Similarity:

The comparison was conducted using four model families, including opus-4-6 (frontier), gpt-4.1 (mature mid-range), gemini 3.1-flash-lite, and the open-weight anchor, DeepSeek v3.2. To guard against prompt artifacts, three independently worded prompt variants per design family were tested, confirming that the results are a property of the design family rather than a single tuned prompt.

Results Across Datasets

The performance comparison showed directional differences based on the benchmark:

(1) TruthfulQA:

The atomic judge shows a small atomic edge, while both judges remain close. The advantage is noted in slightly better unsupported detection.

(2) ASQA and QAMPARI:

In these completeness-heavy settings, the holistic judge is directionally stronger across all four model families—with statistically reliable gains in three of four. The holistic advantage is specifically concentrated in partially supported cases—incompleteness detection. For ASQA, the advantage on partially supported cases ranged from +14.5pp to +33.0pp for proprietary families.

Key Findings and Analysis

The main conclusion is that a matched holistic rubric is competitive with the self-decomposing single-prompt pattern on two of three benchmarks and cheaper on all three. Token accounting revealed that atomic judges produce substantially more tokens (1.4–2.3× the holistic median) on ASQA and QAMPARI, reinforcing that decomposition does not buy accuracy commensurate with its token overhead on completeness-heavy tasks. Furthermore, ablation studies confirmed that increasing the output budget did not rescue the atomic judge's performance gap on ASQA, ruling out truncation as an explanation. The stability across prompt variants confirms the finding is a property of the design family, not a prompt-wording artifact.

Robustness and Sensitivity Checks

The study also examined robustness against perturbations:

(1) Reference Quality Degradation:

Reference-quality degradation produced the largest accuracy drops for both judge families, suggesting neither design provides inherent robustness against incorrect grounding.

(2) Perturbation Stability:

Both judge families showed minimal order sensitivity, with flip rates being low across swapped references, distractors, and verbosity padding. This indicates that the dominant sensitivity is to reference quality, not presentation format. The findings are conditional on the specific self-decomposing single-prompt pattern tested and do not generalize to multi-stage atomic pipelines or non-QA tasks.

Conclusion

The paper concludes that for benchmark-style completeness-sensitive reference support classification, a matched holistic rubric is competitive with the self-decomposing single-prompt pattern on completeness benchmarks (ASQA, QAMPARI), while TruthfulQA shows a small atomic edge. The holistic advantage is localized to incompleteness detection. Practitioners should consider this finding conditional on the specific design pattern tested and recognize that reference quality degradation poses the greatest threat to accuracy for both designs. The work provides a controlled empirical counterexample to the heuristic that decomposition is automatically advantageous. (Word count: 520)


**(Self-Correction Check: The summary adheres strictly to the requested format, uses direct quotes, and avoids external commentary. It focuses on the comparison of design patterns for completeness-sensitive classification.

Improvements for AI systems

Here are the specific improvements an AI system can make based on this research, categorized by application:


)Prompt-Controlled Holistic Judge Implementation (For High-Stakes Classification)

The core improvement is shifting from a single atomic decomposition strategy to a more robust, context-aware holistic scoring mechanism for completeness-sensitive tasks. This system should utilize the prompt structure of Judge Design D.2 or D.6, which explicitly asks the judge to evaluate against a detailed rubric covering correctness, completeness, unsupported detail, and style bias—without forcing an intermediate claim decomposition bottleneck.

  1. The AI system will be engineered to use a prompt that instructs it to perform a multi-dimensional comparison (factual overlap vs. missing information vs. unsupported additions) rather than just checking discrete claims sequentially.

  2. This allows the judge to directly address the benchmark's completeness criterion, which is known to be stricter than general factual correctness standards on benchmarks like ASQA and QAMPARI.

)Improved System Capabilities: Completeness Detection in Reference-Grounded QA

The improved AI system will excel at tasks requiring nuanced understanding of supplied evidence where omissions are penalized. Specifically:

  1. The system can reliably classify a candidate answer as partially supported when it is factually correct but fails to include specific, expected details or items from a comprehensive reference list (e.g., naming all actors in a series, listing all relevant dates/parties).

  2. It will be significantly more robust against prompt-wording artifacts and output budget constraints compared to purely atomic decomposition pipelines, as demonstrated by the stability across three independent prompt variants for both designs.

)Enhanced Robustness Against Grounding Failures

The AI system's performance will be made more reliable when the supplied reference material is degraded or incorrect.

  1. When faced with a swapped or corrupted reference (a common failure mode in real-world RAG pipelines), the holistically-designed judge is expected to maintain higher accuracy than an atomic judge, suggesting a better ability to generalize and assess overall context rather than failing catastrophically on every single claim verification.

  2. The system can be tuned to be less sensitive to minor presentation changes (style bias or verbosity padding) because the holistic rubric explicitly instructs the judge to ignore style unless it introduces factual errors, offering superior resilience in noisy environments.

)Optimized Token Efficiency for Completeness Tasks

For high-volume evaluation where cost is a factor, the system can be optimized to use fewer tokens while maintaining or exceeding performance on completeness-heavy tasks.

  1. The improved system will leverage the holistic judge design, which was shown to be more token-efficient on ASQA and QAMPARI (producing 1.4–2.3x fewer tokens than atomic pipelines) without sacrificing accuracy in detecting incompleteness (partially supported cases).

  2. This allows for more cost-effective evaluation of benchmark adherence without the overhead of generating extensive intermediate claim lists, which proved to be a bottleneck in the atomic design on these specific tasks.

Sources

Related papers