Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data".
Jane: The paper was written by Florian E. Dorner, Vivian Y. Nastl and Moritz Hardt from Max Planck Institute for Intelligent Systems and Tübingen AI Center and ETH Zürich.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, welcome back to the show, everybody. Today we're digging into a paper that's been making the rounds, and the title alone is a bit of a gut punch: "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data."
Jane: Tom, I love this title because it's basically a spoiler. It's telling you the punchline right there in the headline. We're talking about the idea of using a strong language model to grade other models, and this paper says, hold on, there's a hard ceiling on how much that actually helps.
Tom: Exactly. And for our listeners who are just tuning in, the whole "LLM as judge" idea is pretty seductive, right? You've got a super smart model, say GPT-four and you want to evaluate a brand new model. Instead of paying humans to label thousands of examples, you just ask GPT-four to do the grading.
Jane: And it seems to work, at least on the surface. The paper even shows that a judge like GPT-four can have eighty-four percent accuracy on MMLU, which sounds great. But the problem is, that judge brings its own biases, its own blind spots. It might systematically prefer its own style of answer, or it might just be wrong in ways that are correlated with the model it's judging.
Tom: Right, and that's where the "frontier" part of the title comes in. The frontier is when you're trying to evaluate a model that's *better* than the judge. That's the whole point of building new models, right? And this paper says, in that exact scenario, the judge's biases are so problematic that all the clever debiasing in the world can't save you.
Jane: So the promise is that you can use a few expensive, high-quality labels to "correct" the cheap, biased judgments from the LLM. And the paper's main result is that this correction can, at best, double your effective sample size. You can't get a tenfold or a hundredfold improvement.
Tom: A factor of two. That's it. So if you needed a thousand human labels to get a reliable score, using an LLM judge and all the debiasing tricks might get you down to needing five hundred. That's not nothing, but it's a far cry from the "scalable evaluation" dream.
Jane: And it's a very clean, theoretical result, too. It's not just an empirical observation. They prove that this is a fundamental limit for any unbiased estimator that uses this kind of data. That's what makes it so powerful.
Tom: So the dream of just having a superhuman model grade everything for free? This paper is a cold shower on that. And I'm curious to hear what our resident engineer, Meng, thinks about this. Meng, you're the one who has to actually build these evaluation pipelines.
Meng: Honestly, Tom, I'm not surprised. We've seen in practice that you can't just trust a model's self-assessment. We've had to build a lot of guardrails and human checks into our own eval systems. This paper gives a solid theoretical reason for why that's necessary, and it gives us a hard number for the best-case scenario. It's useful for setting expectations with the team.
Tom: So it's not just a theoretical curiosity, it's a practical constraint for people building real systems.
Meng: Exactly. It tells us where to focus our engineering effort. We shouldn't be pouring all our resources into squeezing more out of the LLM judge. We should be figuring out how to get the most information out of our limited ground-truth labels.
Jane: And that's a great segue, because the paper doesn't just say "it's hopeless." It actually gives a framework for thinking about the problem, and it points to where the real gains might be. We'll get into that in the next segment.
Tom: Stay with us, folks. We're just getting started with "Limits to scalable evaluation at the frontier."
Summary: Jane: Welcome back. We're continuing our discussion of "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data." Last segment, we talked about the big, somewhat depressing headline. Now let's get into the actual meat of the paper and how they arrived at that conclusion.
Tom: Yeah, and I think the key is that they set up a very formal, clean problem. They're not just looking at one specific benchmark. They define a general setup where you have a model, a task, and a binary score. Is the answer correct? Is the response safe? Did the model win the arena match?
Jane: And they have a judge model that provides a proxy for that score. The judge is also giving a binary answer. So now you have two random variables for each input: the true score and the proxy score. And the whole problem is about the relationship between those two.
Tom: Right. And the crucial insight is that the value of the proxy score is entirely determined by its correlation with the true score. If the proxy is perfectly correlated with the truth, then it's just as good as the truth. If it's completely uncorrelated, it's useless.
Jane: And their main theorem is about that correlation. They show that if the judge is less accurate than the model being evaluated, then the squared correlation between the proxy and the truth is, at most, zero point five.
Tom: And that zero point five number is the magic one. Because the best possible sample efficiency you can get from any debiasing method is one over one minus that squared correlation. So if the max squared correlation is zero point five, the max sample efficiency is one over zero point five, which is two.
Meng: So the math is pretty elegant. The theoretical limit on the correlation directly translates to a limit on the sample efficiency. It's a really tight chain of reasoning.
Jane: Exactly, Meng. And they're not just hand-waving about the debiasing methods, either. They prove that the method they analyze, which is a form of Prediction Powered Inference, is essentially optimal. So you can't do better with a different, smarter algorithm.
Tom: And they also address a common misconception. People often point to the judge's agreement rate with human experts as proof that it's good. But this paper shows that a high agreement rate, even ninety-nine percent, doesn't guarantee that the judge is useful for ranking models.
Jane: Because the agreement rate can be high while the judge's errors are still systematically biased against better models. It's like a judge in a competition who gets most calls right but consistently penalizes the favorite. The overall accuracy is high, but the ranking is still wrong.
Tom: And they have this nice proposition that shows you can have zero judge bias on average, but still not be able to use the proxy at all. The variance of your estimate doesn't improve at all. So even in the best-case scenario for the judge, the proxy can be worthless.
Meng: That's a brutal result. So it's not just about the average bias. It's about the variance in the estimate of that bias. If you have to estimate the bias for each new model, and that estimation is as hard as the original problem, you're not saving anything.
Jane: Precisely. And that's why the factor-of-two limit is so robust. It's not a bug in their specific method. It's a fundamental property of the information available in the proxy scores when the judge is weaker than the model.
Tom: So the summary is: LLM judges are useful, but only up to a very hard limit. And that limit is a factor of two in sample efficiency. It's a sobering but very well-supported conclusion.
Jane: And it raises a big question: if this is the limit, what are we supposed to do? Where do we go from here? And that's exactly what the paper's experiments and discussion are about. We'll dive into that next.
Tom: Don't go anywhere. We'll be right back.
Improvements: Jane: Welcome back to the show. We're still on "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data." So we've established this hard limit. The question now is, what does the paper suggest we do about it? Is there any way to get around this factor-of-two ceiling?
Tom: Right, and the paper is actually pretty good about this. They don't just leave you with the bad news. They explore a few avenues where the limit might not apply. And the first one is stratification.
Meng: Stratification, meaning you break the problem down into smaller, more homogeneous groups. Instead of evaluating a model on all of MMLU at once, you evaluate it on each subtask separately.
Jane: Exactly. And the paper's experiments on MMLU subtasks are really interesting. They show that within each subtask, the sample efficiency factor can be higher than two. In some cases, it's even above two.
Tom: But here's the catch. The overall sample efficiency is still bounded by the average across all strata. And in their experiments, even with stratification, the overall gains were still modest. They saw some per-stratum gains above two, but they were rare and they didn't translate into a huge overall win.
Meng: So stratification helps, but it doesn't fundamentally break the limit. It just shifts the problem around a bit.
Jane: Right. And the second avenue they explore is making the proxy score continuous instead of binary. So instead of just saying "this answer is correct," the judge gives a probability. Like, "I'm eighty percent sure this answer is correct."
Tom: And that makes intuitive sense. A continuous score carries more information than a binary one. You're not throwing away the judge's uncertainty. And their experiments show that this does help. The sample efficiency factor goes up.
Meng: But I bet it still doesn't break the factor-of-two barrier in the frontier case.
Jane: You'd bet correctly, Meng. Their Theorem ten shows that even with a continuous proxy, if the judge is still worse than the evaluated model, you're still stuck with a maximum sample efficiency factor of two. The continuous score helps, but it doesn't change the fundamental limit.
Tom: So the paper is really hammering home the point that the limit is about the relative capability of the judge and the evaluated model, not about the specific form of the proxy.
Jane: And that's the key takeaway. The only way to get around the limit is to make the evaluation task easier than the model's task. If judging is fundamentally easier than doing, then you can get gains. But at the frontier, that's rarely the case.
Meng: So what's the practical advice for someone like me? We're building a system to evaluate a new model that we think is state-of-the-art. What do we do?
Tom: Well, the paper's advice is pretty clear. Don't expect the LLM judge to save you. You still need to invest in getting high-quality ground-truth labels. The LLM judge can give you at most a factor of two, and in practice, it's often less.
Jane: And they also point out that this applies to any biased evaluator, not just LLMs. If you're using crowdworkers who aren't experts, you're in the same boat. The limit is about the quality of the proxy relative to the truth, not about the source of the proxy.
Meng: That's a really important point. It means we need to think carefully about where we get our labels, not just how we process them.
Jane: And it also points to a promising avenue for future work. If we can build specialized evaluators that are genuinely better at the evaluation task than the model is at the original task, then we can get real gains. But that's a much harder problem than just prompting a general-purpose LLM.
Tom: So the improvements suggested are really about changing the nature of the evaluation task itself, not just tweaking the algorithm. That's a much deeper insight. And it sets us up perfectly for our final segment, where we'll wrap up and think about the bigger picture.
Conclusion: Tom: And we're back for the final stretch. We've been talking about "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data." And I think it's time to step back and think about what this all means for the field.
Jane: For me, the biggest takeaway is that the promise of fully automated, scalable evaluation is a mirage, at least for the models we care about most. The frontier is exactly where you need the most careful evaluation, and that's exactly where LLM judges are the least useful.
Tom: And it's not because the judges are bad. It's because the information they provide is fundamentally limited when they're not better than the thing they're judging. The math is really clear on that.
Meng: I think the practical impact is huge. It means that a lot of the hype around "self-evaluating" models is misplaced. You can't just have a model grade itself and expect to get reliable results. You need a source of truth that's independent and of high quality.
Jane: And that source of truth is still going to be human experts, at least for the foreseeable future. The paper doesn't say that human evaluation is easy. It just says that you can't avoid it if you want to know how good your frontier models really are.
Tom: And that has implications for how we allocate resources. Instead of spending all our time building clever debiasing algorithms, maybe we should be investing in better annotation tools, better ways to train human experts, and better ways to sample the data that we do label.
Meng: I agree. And I think the paper also gives us a useful framework for thinking about when LLM judges *are* useful. If you're evaluating a model that's clearly weaker than the judge, then the gains can be real. So there's still a place for this technology, just not at the very edge of capability.
Jane: And that's a nice, nuanced conclusion. It's not "LLM judges are useless." It's "LLM judges are useful, but only within their limits." And this paper gives us a precise, mathematical way to understand those limits.
Tom: It's a really important paper for anyone working on model evaluation, whether you're a researcher, an engineer, or just someone who cares about how we measure progress in AI. It's a reality check, but a very valuable one.
Jane: So we'll say goodbye to "Limits to scalable evaluation at the frontier." It's been a great discussion. And we're ready to move on to the next paper.
Tom: Thanks for listening, everyone. We'll see you next time.
Florian E. Dorner, Vivian Y. Nastl, Moritz Hardt
Max Planck Institute for Intelligent Systems · Tübingen AI Center · ETH Zürich
cs.LG, stat.ML
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: ICLR 2025; 28 pages, 8 figures
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 71/100
Key concepts
- LLM as Judge
- This method of using a highly capable language model to act as an evaluator or grader for other AI models. Instead of relying on human experts, the LLM provides a proxy score for correctness or quality.
- Sample Efficiency
- This measures how much data is required to get a reliable evaluation score. The paper proves that even with debiasing techniques, the maximum improvement in sample efficiency is limited to a factor of two when using an LLM judge.
- The Frontier
- This describes the challenging scenario when evaluating a new AI model that is potentially superior or 'state-of-the-art' compared to the language model being used as the judge. This is where the fundamental limitations of the LLM judge become most evident.
Terminology
Summary
Affiliation: Max Planck Institute for Intelligent Systems, Tübingen; Tübingen AI Center; ETH Zürich
Published: arXiv:2410.13341v3 [cs.LG] 6 Jan 2026
"High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an important research ambition. Many hope to use strong existing models in lieu of costly labels to provide cheap model evaluations. Unfortunately, this method of using models as judges introduces biases, such as self-preferencing, that can distort model comparisons. An emerging family of debiasing tools promises to fix these issues by using a few high quality labels to debias a large number of model judgments. In this paper, we study how far such debiasing methods, in principle, can go. Our main result shows that when the judge is no more accurate than the evaluated model, no debiasing method can decrease the required amount of ground truth labels by more than half. Our result speaks to the severe limitations of the LLM-as-a-judge paradigm at the evaluation frontier where the goal is to assess newly released models that are possibly better than the judge. Through an empirical evaluation, we demonstrate that the sample size savings achievable in practice are even more modest than what our theoretical limit suggests. Along the way, our work provides new observations about debiasing methods for model evaluation, and points out promising avenues for future work."
The paper addresses the challenge of evaluating newly released AI models as they advance in capabilities. Expert data annotation is not only slow and costly. Traditional benchmarking also struggles to keep up with rapidly changing model capabilities across an expanding range of tasks.
The authors examine the model-as-judge
paradigm, where a strong existing model to provide judgments about other models
is used to replace human annotators.
However, the authors note that When used as judges, models exhibit a range of biases that can skew model comparisons and result in misleading model rankings.
Recently proposed debiasing methods promise a compelling way forward. Using a small number of ground truth labels, these methods can potentially debias a large number of model predictions, thus restoring their utility for benchmarking purposes.
The paper focuses on the evaluation frontier: Newly released models for which we have little intuition as of yet. What makes this case so challenging is that the new model is likely better than the judge in some ways.
The central question is whether debiasing methods together with the model-as-judge paradigm can, in principle, provide an adequate solution to scalable evaluation at the frontier.
The main prediction is sobering: "Whenever the judge model performs worse at its task than the evaluated model, the optimal debiasing method is no better than using twice the ground truth data. This shows that, although there is merit to debiasing, at the evaluation frontier its economic gains are not greater than a factor-two savings in annotation cost."
The paper analyzes aggregated binary evaluations: Given a prompt x, a model m receives a binary score s(x, m(x)) depending on its output m(x).
The model is evaluated based on its expected score E s(X, m(X))
where the expectation is taken over the prompt distribution X ∼ D, as well as additional randomness in the scores s(·, m(·)) introduced by the evaluation protocol.
This setup covers:
-
Accuracy in classification and Q&A benchmarks:
the accuracy score s(x, m) equals one whenever m(x) equals the label y(x)
-
Arena-style benchmarks:
The score s(x, m) indicates whether m(x) was judged to be a better response than m′(x)
-
Safety benchmarks:
The score s(x, m) indicates whether the model's response m(x) is safe
A judge model provides a proxy score s̃
that is cheaper to obtain. The joint distribution of the two induced random variables s(m) and s̃(m) is specified by three parameters:
-
b(m) = P s(m) = 1 — the model's expected real score
-
p(m) = P s̃(m) = s(m) s(m) = 1 — the true positive rate
-
q(m) = P s̃(m) = s(m) s(m) = 0 — the true negative rate
The paper defines judge bias as JB(m):= E(s̃(m) − s(m)) = (1 − q(m))(1 − b(m)) − (1 − p(m))b(m). Whenever JB(m) has large magnitude, the proxy score s̃ misrepresents the performance b(m) of model m, even with large sample sizes.
Proposition 1 shows that when using a classifier m̃ to evaluate a set M of strictly better classifiers (where m̃(x) = y(x) implies mi(x) = y(x) for all mi ∈ M), "E s(mi) > E s(mj) implies E s̃(mi) < E s̃(mj)" — meaning the proxy evaluations fully reverse the correct model ranking.
The paper defines agreement as AG(m):= P s(m) = s̃(m) = b(m)p(m) + (1 − b(m))q(m).
Proposition 2 shows that for agreement AG(m) = r, there can be a judge bias of 1 − r in either direction.
This means even if AG(m) = r was constant across models m, we could only reliably rank models for which the true score E s(m) differs by more than 2(1 − r).
The authors conclude: without further assumptions an agreement rate above 99% (for each evaluated model!) would be required to ensure (asymptotically) correct rankings.
Proposition 3 states that if all model evaluations are approximately unbiased (E θ̂(m) − E(s(m)) < ϵ/2 for models separated by at least ϵ), then using θ̂ to rank models yields the correct ranking with high probability
as variance converges to zero.
The paper follows the PPI framework (Angelopoulos et al., 2023a). The PPI estimator is:
θ̂ PP(x) = (1/N) Σ i=n+1 N+n s̃ i(m) + (1/n) Σ j=1 n (s j(m) − s̃ j(m))
This estimator is unbiased: E θ̂ PP = E s̃(m) + E s(m) − E s̃(m) = E s(m).
The variance is: Var θ̂ PP = (1/N) Var s̃(m) + (1/n) Var(s̃(m) − s(m)).
An interpolated version θ̂ λ PP = λθ̂ PP + (1−λ)θ̂ GT with optimal λ* = Cov(s(m), s̃(m))/(1 + n/N) Var s̃(m) never increases variance compared to the classical ground truth estimator.
The paper defines the sample efficiency factor as τ(θ̂):= Var θ̂ GT / Var θ̂ = 1/r, where r = Var θ̂ / Var θ̂ GT.
Proposition 4 provides an upper bound: τ(θ̂ λ PP*) ≤ 1/(1 − ρ(s(m), s̃(m))2), where ρ2 is the squared Pearson correlation between s and s̃.
Theorem 5 states: "Let Θ be the set of all unbiased estimators for E s(m) that observe n joint IID samples (s(m), s̃(m)) and N IID proxy samples s̃(m) independent of the joint samples. Then, any θ̂ ∈ Θ fulfills the variance bound Var θ̂ ≥ Var θ̂ λ PP*."
This means the sample efficiency factor of θ̂ is bounded by τ(θ̂) ≤ max θ̂∈Θ τ(θ̂) = τ(θ̂ λ PP*) ≤ 1/(1 − ρ(s(m), s̃(m))2).
Theorem 6: Assume 0.5 ≤ AG(m) ≤ b(m). Then, ρ(s(m), s̃(m))2 ≤ 0.5.
Corollary 7: Assume 0.5 ≤ AG(m) ≤ b(m). Then, τ max = max θ̂∈Θ τ(θ̂) ≤ 2.
The authors explain: "Corollary 7 implies that when evaluating state-of-the-art models, the best we can expect from using LLM judges is a factor-two improvement in sample efficiency. This is unless the judge's task is substantially easier than the evaluated model's task."
Proposition 8: "For any agreement rate 0.5 ≤ AG(m) < 1, there exist values of b(m), p(m), q(m) ∈ (0, 1) such that JB(m) = ρ(s(m), s̃(m))2 = 0 and thus τ max = 1."
This shows that in some cases despite access to s̃, no unbiased estimator has less variance than θ̂ GT.
The paper evaluates models on MMLU (Hendrycks et al., 2021) using predictions from the HELM leaderboard. Figure 2 shows that Sample efficiency gains stay below two, unless SOTA models are used to evaluate weak models.
Specifically, τ max consistently stays below the value of two suggested by Corollary 7, except when current flagship models are used to judge the significantly worse LLama2-7b on MMLU.
For MT-Bench (Zheng et al., 2024), Figure 3 shows that Sample efficiency get close to two in some cases, but consistently stay below that value.
The authors note: "Interestingly, τ max stays below two in all other cases, even when stronger models like GPT-4 are used to evaluate weaker models like LLama3-70B. This suggests that the upper bound on τ max from Corollary 7 is fairly robust."
The paper relaxes the assumption of binary proxy scores, allowing continuous proxies such as the probability a judge assigns to a model's answer.
Proposition 9: For any binary proxy s̃ such that 0.5 ≤ AG ≤ b, we have 0.5 ≤ SO(R(s̃)) ≤ b,
where SO is the soft agreement and R(s̃) is the recalibrated proxy.
Theorem 10: For any proxy score s̃ with SO(R(s̃)) ≤ b, we have ρ2(s, s̃) ≤ 0.5. Correspondingly, the sample efficiency of PPI is bounded: τ(θ̂ λ PP*) ≤ 2.
Experiments using LLama3.1-405B's next-token predictions on MMLU show that "using the non-binary proxy consistently improves sample efficiency. However, as suggested by Theorem 10, the sample efficiency factor τ(θ̂ λ PP*) remains below two when we use LLama3.1-405B to evaluate the stronger Claude 3.5."
Proposition 11: For strictly better classifiers with q(m) = 0, b(m) = x + δ, and p(m) = x/(x+δ), ρ(s(m), s̃(m))2 ≤ 1/9. Correspondingly, the sample efficiency factor is bounded by τ max ≤ 1.125.
Theorem 12: For any value of BA(m), we have: 4b(m)(1−b(m))(2 BA(m) − 1)2 ≤ ρ(s(m), s̃(m))2 ≤ 2 BA(m) − 1,
where BA is the balanced agreement.
Corollary 13: Whenever BA(m) ≥ 0.5, we have ρ(s(m), s̃(m))2 ≤ min p(m), q(m).
The authors conclude: "Our results show that for evaluating frontier models, LLM judges might fall short of the promise of largely replacing expert labelers: While doubling the effective sample size can be useful for practitioners, the order of magnitude of required ground truth labels remains the same with and without access to LLM judges."
They note three ways to potentially circumvent the negative results:
-
Approaches like stratified PPI (Fisch et al., 2024) might be able to obtain a somewhat better sample efficiency factor
— thoughTheorem 5 still applies per-stratum
and empirical resultssuggest that at the frontier, stratified PPI rarely improves sample efficiency by more than a factor of two.
-
"Theorem 6 only holds if the proxy s̃ is less capable at predicting s than the evaluated model m is at its task. Thus, for tasks in which evaluation is substantially easier than the task itself... gains in sample efficiency of more than a factor two are possible."
-
Our results do not only apply to LLM judges, but any form of biased evaluators including (poorly instructed) crowdworkers.
The paper situates itself within the LLM-as-a-Judge literature, noting that LLM judges can be biased in numerous ways
including self-preferencing (Liu et al., 2023; Panickssery et al., 2024), preference for longer outputs (Dubois et al., 2024b; Wei et al., 2024), and choice-order bias (Dominguez-Olmedo et al., 2023; Wang et al., 2023; Shi et al., 2024).
Regarding debiasing methods, the authors note that prior work (Chaganty et al., 2018) found that their method, which is essentially equivalent to PPI, only improved data efficiency by around 10% using 2018's automated metrics.
More recent works (Boyeau et al., 2024; Chatzi et al., 2024; Fisch et al., 2024) consistently show that PPI improves efficiency in terms of ground truth labels, but gains in effective sample size rarely exceed 50% and are always below 100%.
Improvements for AI systems
Based on the paper, here are specific improvements for AI systems:
1. Implement Bias-Corrected Evaluation for Model Comparisons
-
Improvement: When evaluating a new model against a judge model (e.g., using GPT-4 to judge Claude), do not rely solely on the judge’s proxy scores. Instead, collect a small set of ground-truth labels (n) and apply the Prediction Powered Inference (PPI) estimator:
θ̂ PP = (1/N)Σs̃ i + (1/n)Σ(s j - s̃ j). Then optimize the interpolation parameter λ to minimize variance. -
What it can do: Provides unbiased performance estimates, preventing misleading rankings (e.g., avoiding the full reversal shown in Proposition 1). This is critical when the evaluated model may outperform the judge.
2. Enforce a Sample Efficiency Ceiling of 2× at the Frontier
-
Improvement: Before deploying an LLM-as-a-judge for a new model, check if
AG(m) ≤ b(m)(judge agreement ≤ evaluated model’s true score). If true, cap the expected benefit of debiasing at a factor of 2 in ground-truth sample efficiency (Corollary 7). Do not allocate resources expecting more than a 2× reduction in annotation cost. -
What it can do: Prevents overinvestment in complex debiasing pipelines for frontier models where gains are fundamentally limited. Guides researchers to instead double the ground-truth data collection if more accuracy is needed.
3. Use Non-Binary Proxy Scores with Recalibration for Slightly Better Efficiency
-
Improvement: Instead of forcing the judge to output a single binary label, extract the judge’s probability distribution over answers (e.g., next-token probabilities for MMLU options). Recalibrate this proxy
R(s̃) = P(s=1s̃)and use it in the PPI estimator. -
What it can do: Achieves sample efficiency factors closer to 2 (and slightly exceeds it for models 5% weaker than the judge), as shown in Figure 4. This is a practical gain over binary proxies, especially when evaluating models slightly below the judge’s capability.
4. Detect and Flag High-Agreement-but-Biased Judges
-
Improvement: Do not trust high agreement rates (e.g., 85% with experts) as a proxy for evaluation quality. Instead, compute the balanced agreement
BA(m) = (p(m) + q(m))/2. IfBA(m)is close to 0.5, the judge provides almost no debiasing value (Theorem 12), even if overall agreement is high. -
What it can do: Identifies judges that appear accurate but are actually useless for ranking (e.g., a judge that always predicts the majority class). Prevents false confidence in evaluation results and prompts the use of ground-truth labels instead.
5. Optimize for Pairwise Ranking Differences, Not Individual Scores
-
Improvement: When the goal is ranking models (not absolute scores), jointly optimize the PPI interpolation parameters λ and λ′ for pairs of models (m, m′) to minimize the variance of the difference
θ̂(m) - θ̂(m′), rather than optimizing each model’s estimate independently. -
What it can do: Can yield better ranking accuracy for the same annotation budget, especially when proxy scores for different models are correlated. This is a direct improvement over standard per-model debiasing.
6. Implement Stratified PPI with Per-Stratum Ceiling Checks
-
Improvement: When using stratified sampling (e.g., by MMLU subtask), apply PPI per stratum, but first verify that the judge is not better than the evaluated model in each stratum. If the judge is better in a stratum, the sample efficiency gain can exceed 2× there; otherwise, cap it.
-
What it can do: Maximizes debiasing gains where possible (e.g., in subtasks where the judge excels) while avoiding wasted effort in strata where gains are capped. Empirically, this rarely exceeds 2× at the frontier (Appendix B.3), but it can be exploited when the judge is stronger on specific domains.
7. Avoid Using Judge Scores for Models That Are Strictly Better
-
Improvement: If you know a model m is strictly better than the judge (e.g., the judge’s correct predictions are a subset of m’s correct predictions), do not use the judge’s scores even with debiasing. Proposition 11 shows the maximum sample efficiency gain is only 1.125× in this case.
-
What it can do: Saves annotation budget by directly using ground-truth labels instead of a debiasing pipeline that offers negligible benefit. This is a clear operational rule for practitioners.
Abstract
High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an important research ambition. Many hope to use strong existing models in lieu of costly labels to provide cheap model evaluations. Unfortunately, this method of using models as judges introduces biases, such as self-preferencing, that can distort model comparisons. An emerging family of debiasing tools promises to fix these issues by using a few high quality labels to debias a large number of model judgments. In this paper, we study how far such debiasing methods, in principle, can go. Our main result shows that when the judge is no more accurate than the evaluated model, no debiasing method can decrease the required amount of ground truth labels by more than half. Our result speaks to the severe limitations of the LLM-as-a-judge paradigm at the evaluation frontier where the goal is to assess newly released models that are possibly better than the judge. Through an empirical evaluation, we demonstrate that the sample size savings achievable in practice are even more modest than what our theoretical limit suggests. Along the way, our work provides new observations about debiasing methods for model evaluation, and points out promising avenues for future work.
Sources
- GPT-4 Technical Report
- PPI++: Efficient Prediction-Powered Inference
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- AutoEval Done Right: Using Synthetic Data for Model Evaluation
- The price of debiasing automatic metrics in natural language evaluation
- Prediction-Powered Ranking of Large Language Models
- Can Large Language Models Be an Alternative to Human Evaluations?
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Questioning the Survey Responses of Large Language Models
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Stratified Prediction-Powered Inference for Hybrid Language Model Evaluation
- GPTScore: Evaluate as You Desire
- Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
- Benchmarking Cognitive Biases in Large Language Models as Evaluators
- Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
- Why do classifier accuracies show linear trends under distribution shift?
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks