No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

summary

Video file (mp4)

The gist

The paper introduces the Business and Finance Fundamentals Benchmark (BFF-Bench), a dataset of 160 challenging questions and long-form responses authored by financial professionals, and the Verified

In short

The hosts discuss a paper titled "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding." They conclude that using one AI model to judge another is unreliable for correctness. Reliability depends on the judge's own competence or providing a verified, correct reference, requiring human verification in an evaluation pipeline.

Key concepts

LLM Judge
An AI model used to evaluate or grade responses from other LLMs. The paper shows that these judges are unreliable for determining factual correctness unless they have a correct answer to compare against.
BFF-Bench
A new benchmark created by the authors, it consists of 160 challenging, multi-turn questions about business and finance. It is used to test whether LLM Judges can accurately assess correctness in high-stakes domains.
Human Grounding
The process of using human expertise to provide verified correct answers or references. The paper argues that this is necessary for an LLM Judge to avoid errors, as the judge relies on accurate input.

Terminology used across episodes

This episode discusses

The paper

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding · Read on arXiv

Kensho Technologies · MIT

Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is appealing due to its scalability, low cost, and strong correlations with human stylistic preferences. However, it remains unclear how accurately these methods can assess response quality in domains where correctness matters more than style. To address this gap, we introduce the Business and Finance Fundamentals Benchmark (BFF-Bench), a dataset of 160 challenging questions and long-form responses authored by financial professionals. These experts subsequently evaluated the correctness of 1,200 responses generated by a diverse set of LLMs on both BFF-Bench and a challenging subset of MT-Bench. With this expert-annotated dataset of judgments (VERDICTS), we analyze the agreement between a suite of automated grading methods and human experts. While we observe that LLM Judges are more reliable than other grading methods, our findings reveal a clear pattern in LLM Judge performance: when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves. We demonstrate that providing the judges with expert-written references largely mitigates this issue, highlighting the limits of using LLM-as-a-Judge without any form of human verification.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding".

Jane: The paper was written by Michael Krumdick, Charles Lovering, Varshini Reddy, Seth Ebner and Chris Tanner from Kensho Technologies and MIT.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re cracking open a paper that’s been making waves in the AI evaluation world, and it’s called "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding." Jane, I have to say, that title alone got me hooked.

Jane: It really grabs you, doesn’t it, Tom? The phrase "No Free Labels" is perfect. It’s telling us that when we use one AI to grade another AI, we can’t just assume the grades are trustworthy without some human input somewhere in the loop.

Tom: Right, and the authors are from Kensho Technologies and MIT. They’ve put together this massive study to test whether these so-called "LLM Judges" are actually any good at judging correctness, not just style.

Jane: And that’s the crucial distinction. A lot of the popular benchmarks out there, they’re measuring things like how engaging or fluent a response is. But in fields like business and finance, getting the facts right is the whole ballgame. So they built a new benchmark called BFF-Bench to test exactly that.

Tom: BFF-Bench. I love it. It’s a set of one hundred sixty really tough, multi-turn questions about business and finance, written by actual financial professionals. And they didn’t just write the questions, they wrote the gold-standard answers too.

Jane: Exactly. And then they took those questions and they had six different AI models generate responses. That gave them one thousand two hundred responses to grade. Then, the same financial experts went through every single one of those responses and judged them as correct or incorrect.

Tom: So we have a mountain of data where we know the ground truth. Now, the question is, when you plug in an LLM Judge to do the same grading, does it agree with the experts?

Jane: That’s the heart of it. And the title is a warning that the answer is often no, especially when the judge isn’t given a correct answer to compare against. It suggests there’s no shortcut to getting those human labels, even if it’s just to verify the references.

Tom: So, no free lunch in the AI evaluation world. You either need a judge that already knows the answer, or you need a human to provide a correct reference. We’re going to dig into exactly how they figured that out next. Stick around.

Summary: Tom: So, Jane, we just set the stage with this paper, "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding." Let’s get into the meat of it. What did they actually do?

Jane: Well, they ran a bunch of experiments. They took these LLM Judges and gave them different types of help. In one case, they gave them no reference at all. In another, they let the judge generate its own reference answer. And in the final case, they gave the judge the human-written, verified correct answer.

Tom: And the results were pretty stark. When the judge was given the human-written reference, its agreement with the human experts shot way up. It was the best performance across the board, for every single judge model they tested.

Jane: Right. But the most fascinating finding was about the judges themselves. They split the questions into two groups: ones where the judge model, acting as a candidate, could answer correctly, and ones where it couldn't.

Tom: And that’s where the title really hits home. When the judge was evaluating responses to questions it already knew how to answer, it did pretty well, even without a reference. But on questions it couldn't answer itself, its agreement with the humans plummeted.

Jane: It’s like asking me to grade a calculus exam. I might be able to tell a good essay from a bad one, but I have no idea if the math is right. These LLM Judges are the same. If they don’t know the answer, they’re just guessing.

Tom: And the paper shows that guessing is not good enough. They even tried giving the judges a "wrong" reference, one that had been subtly edited to be incorrect. And you know what happened? The judges’ performance got even worse than when they had no reference at all.

Jane: That’s the scary part. A slightly wrong reference is more dangerous than no reference, because the judge trusts it and gets led astray. They also tested random references from other questions, which didn't help either.

Tom: So the summary is clear. The reliability of an LLM Judge hinges on either its own competence on the topic or the correctness of the reference it’s given. There’s no way around it.

Jane: And that’s why they call it "No Free Labels." You can’t just automate the whole evaluation pipeline and expect to get trustworthy results. You need that human grounding, at least to verify the references. Now, let’s talk about what this means for the way we should actually build these evaluation systems.

Improvements: Tom: Welcome back. We’re deep in "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding," and we’ve seen the problem. So, Jane, what’s the fix? What are the authors suggesting we do differently?

Jane: The biggest suggestion is that we need to stop treating LLM Judges as a complete replacement for human evaluation. Instead, we should use them in a way that acknowledges their limitations. The paper’s main recommendation is to provide the judge with a correct, verified reference.

Tom: And they even found a way to make that more scalable. They compared using the human-written references against references that were generated by a strong model like GPT-4o but then verified by humans as correct. The performance was nearly identical.

Jane: That’s a huge finding. It means you don’t necessarily have to write every reference from scratch. You can have a model draft it, and then a human just checks it for accuracy. That’s a much more efficient use of human expertise.

Tom: So it’s not about eliminating humans, it’s about moving them to a higher-level task. Instead of grading one thousand two hundred responses, they just need to verify one hundred sixty references.

Jane: Exactly. And the paper is very clear about the dangers of not doing this. They showed that using a judge’s own generated reference, which is unverified, actually makes the self-preference bias worse. The judge gets even more likely to favor its own responses.

Tom: So the "Self" reference wasn't just unhelpful, it was actively harmful in some cases. That really drives home the point that the correctness of the reference is the most important thing, not just having a reference at all.

Jane: Right. The authors are essentially saying, "If you’re going to use LLM-as-a-Judge, you must ground it in verified facts." It’s a call for more rigorous evaluation practices, especially in high-stakes domains like finance.

Tom: And it gives us a practical roadmap. Use a strong model to draft references, have a human verify them, and then let the LLM Judge loose. That’s a workflow that can actually scale. Now, let’s look at the very first page of the paper to see how they framed this whole problem.

First Page: Tom: So, Jane, we’ve talked about the findings and the recommendations. Let’s go back to the very beginning of "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding" and see how they set the stage.

Jane: The abstract immediately frames the problem. They point out that LLM-as-a-Judge is popular because it’s scalable, cheap, and matches human *stylistic* preferences well. But they ask a critical question: can it judge *correctness*?

Tom: And that’s the gap they’re trying to fill. They introduce BFF-Bench, which we talked about, and they also mention they created a corrected version of a subset of MT-Bench, which is a popular existing benchmark.

Jane: That was a really interesting move. They found that fifteen out of forty of the reference answers in MT-Bench’s math and reasoning category were actually incorrect or inconsistent. So they had to fix them before they could use them.

Tom: So they’re not just creating new data, they’re also cleaning up existing data. That shows how pervasive this problem of unreliable references is.

Jane: Exactly. And then they introduce VERDICTS, which is the dataset of all those expert judgments on the one thousand two hundred responses. It’s the ground truth that makes their whole analysis possible.

Tom: The first page really lays out the core promise. They’re going to show that an LLM Judge’s agreement with humans is tied to its own ability to answer the question, and that providing a correct reference largely fixes the problem.

Jane: And they put that key figure right there on the first page. It shows the agreement dropping off a cliff for questions the judge gets wrong, and then recovering when you give it the human reference.

Tom: It’s a really compelling visual. It’s like they’re showing you the whole story in one picture. The paper is saying, "Here’s the problem, and here’s the data to prove it."

Jane: And it sets up the rest of the paper perfectly. They’re going to rigorously test this hypothesis across different models, different tasks, and different reference types. It’s a very thorough investigation.

Tom: So from the very first page, they’re making a strong case that we need to be much more careful about how we evaluate our AI models. We’re going to wrap this up with our final thoughts next.

Conclusion: Tom: Alright, we’ve reached the end of our journey with "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding." Jane, can you give us the big picture one more time?

Jane: Absolutely, Tom. The paper’s core message is that using an LLM to judge another LLM is not a reliable shortcut for determining correctness. It works well when the judge already knows the answer, or when it’s given a verified, correct reference to compare against.

Tom: And the moment you take away that support, the whole thing falls apart. The judges start agreeing with humans on questions they can answer, and basically guessing on everything else.

Jane: Right. And we learned that a wrong reference is worse than no reference at all. It actively misleads the judge. So the quality of the reference is everything.

Tom: So the call to action is pretty clear. If you’re building an evaluation pipeline, you need to invest in human verification. It’s the only way to ensure your benchmarks are actually measuring what they claim to measure.

Jane: And that’s the "No Free Labels" part. You can’t get away from the need for some human expertise. But the paper shows a smart way to use it, by having humans verify model-generated references instead of writing everything from scratch.

Tom: It’s a sobering but really important message for the field. We’re building these massive models, and we need to make sure we’re evaluating them honestly.

Jane: Well said, Tom. It’s a fantastic piece of research, and we’re sad to see it go. But we’re excited to see what’s next in the world of AI evaluation.

Tom: Thanks for listening, everyone. We’ll be back soon with another paper to break down. Until then, keep questioning the results you see.

More episodes

← Home