On Benchmarking Human-like Intelligence in Machines

summary

Video file (mp4)

The gist

This paper argues that current evaluation paradigms for AI are insufficient for assessing human-like cognitive capabilities, and proposes recommendations for future benchmarks based on cognitive

In short

The episode discusses a paper titled "On Benchmarking Human-like Intelligence in Machines." The hosts explain how current AI benchmarks often fail to match human judgment, finding that labels are often flawed or ambiguous. They conclude that moving beyond single 'correct' answers and focusing on collecting real human data is necessary for building trustworthy, human-like AI.

Key concepts

Human Evaluation Study
The researchers tested nine popular AI benchmarks by having over 100 people rate the same questions. Instead of picking one answer, participants used sliders to express graded agreement and also indicated how confident they were in their response.
Structured Disagreement
This occurs when a task is genuinely ambiguous, and different people hold opposite views with high confidence. The benchmark label simply picks one answer as 'truth,' ignoring this natural variation in human judgment.
Gradedness and Uncertainty
This involves measuring not just what people think, but how strongly they believe it. People often give intermediate slider values with high confidence, indicating the true answer lies somewhere in the middle.

Terminology used across episodes

This episode discusses

The paper

On Benchmarking Human-Like Intelligence in Machines · Read on arXiv

Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum

Harvard University · University of Cambridge · Massachusetts Institute of Technology · New York University · Johns Hopkins University · Princeton University · Stanford University

Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition. However, we argue that many current evaluation paradigms for AI are insufficient for assessing human-like cognitive capabilities in these models. We identify a set of key shortcomings: a lack of human-validated labels, inadequate representation of human response variability and uncertainty, and reliance on simplified and ecologically invalid tasks. We support our claims by conducting a human evaluation study on nine existing AI benchmarks, suggesting major limitations in task and label designs. To address these limitations, we propose five concrete recommendations for future AI evaluation efforts that will enable more rigorous and meaningful understanding of human-like cognitive capacities in AI models.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "On Benchmarking Human-like Intelligence in Machines".

Jane: The paper was written by Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu et al. from Harvard University and University of Cambridge and Massachusetts Institute of Technology and New York University and Johns Hopkins University and Princeton University and Stanford University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we’re digging into a paper that’s got a title that really makes you stop and think: “On Benchmarking Human-like Intelligence in Machines.” Jane, what’s your first reaction to that title?

Jane: Honestly, Tom, it feels like a bit of a challenge. Like, we keep hearing that AI is getting more human-like every week, but this paper is basically saying, “Hold on, are we even measuring that the right way?” And I love that they’re asking that question.

Tom: Right? And the authors are a who’s who of cognitive science and AI. You’ve got Joshua Tenenbaum and Thomas Griffiths on there, along with a bunch of folks from MIT, Harvard, Princeton. This is serious brainpower.

Jane: Yeah, and that’s what makes it exciting. They’re not just armchair philosophers. They went out and actually tested some of the most popular AI benchmarks with real human participants to see if the labels even make sense.

Lu: That’s the part that got me hooked. They took nine benchmarks—things like moral reasoning, social IQ, theory of mind—and they had over a hundred people rate the same questions. And what they found is that the “correct” answers in these benchmarks often don’t match what real people think.

Tom: Lu, you’re jumping ahead, but I love it. So the title is really a promise. It’s saying, “If you want to claim your AI is human-like, you need to actually compare it to humans, not just to a dataset label.”

Jane: And that’s the big implication, right? We’re building these systems to work with us, to understand us. If the benchmarks we use to test them are flawed, we’re basically flying blind.

Meng: From an engineering standpoint, that’s terrifying. I mean, we tune models to score well on these benchmarks. If the benchmarks are measuring the wrong thing, we’re just optimizing for the wrong target.

Tom: Exactly, Meng. And that’s what we’re going to unpack today. We’re going to look at what they found when they put these benchmarks under the microscope, and then we’ll get into their recommendations for how to fix it.

Jane: And trust me, some of those findings are wild. We’re talking about tasks where human agreement with the benchmark label was below chance. That’s not a small problem.

Tom: So stick around. We’re about to find out why your favorite AI might be acing tests that humans would fail.

Summary: Jane: So, Tom, we’ve set the stage. Now let’s get into the meat of “On Benchmarking Human-like Intelligence in Machines.” What did they actually do?

Tom: They ran a human evaluation study on nine AI benchmarks. They took tasks like irony detection, dark humor, moral permissibility, social support, and they had one hundred seventeen people rate them using sliders instead of multiple choice.

Lu: And that’s a key detail. Instead of forcing people to pick one answer, they let them express graded agreement. So if you think something is *mostly* supportive but not totally, you can put the slider at seventy instead of being forced to say yes or no.

Jane: And they also asked people how confident they were in each answer. So you get two separate measures: what you believe, and how sure you are about that belief.

Tom: Right. And the results are pretty damning. On average, human agreement with the benchmark labels was only about seventy percent. But here’s the kicker: on the Social Support task, humans agreed with the label less than half the time. That’s worse than flipping a coin.

Meng: Wait, so the benchmark label is basically wrong more often than not? How does that even happen?

Jane: That’s the thing, Meng. These labels were probably created by one or two annotators, or maybe by the benchmark designers themselves. They assumed there was one right answer, but when you ask a hundred people, you get a whole distribution of opinions.

Lu: And it’s not just noise. They show examples where people split into two camps, both very confident, but with opposite answers. Like, one person is sure the statement is supportive, another is sure it’s not, and they’re both certain they’re right.

Tom: That’s what they call structured disagreement. It’s not that people are confused; it’s that the task is genuinely ambiguous, and the benchmark just picked one answer and called it the truth.

Jane: And that’s a huge problem if you’re trying to build a model that behaves like a human. If you train it to match that single label, you’re training it to ignore the natural variation that exists in human judgment.

Meng: So what does that mean for me? If I’m deploying a model that was trained on these benchmarks, am I just shipping something that’s fundamentally misaligned with what people actually think?

Tom: That’s exactly the concern. And it’s not just about accuracy. It’s about whether the model’s *patterns* of judgment—its uncertainty, its errors—look anything like a human’s.

Jane: And that’s the core of their argument. They’re saying we need to stop measuring against a single “correct” answer and start measuring against the full distribution of human responses.

Lu: Which brings us to their recommendations. And that’s where the real actionable stuff is.

Tom: Exactly. So we’ve seen the problem. Now let’s talk about how they think we should fix it.

Improvements: Jane: Alright, so we’ve seen that the current benchmarks are shaky. What does “On Benchmarking Human-like Intelligence in Machines” actually propose we do differently?

Tom: They give us five concrete recommendations, and the first one is almost obvious: actually collect human data. Don’t just make up labels. If you want to claim your AI is human-like, you need real human responses as the gold standard.

Lu: And not just a handful. They’re saying you need robust sample sizes. Because if you ask five people, you might get a very different distribution than if you ask a hundred.

Meng: So that’s recommendation one. What’s next?

Jane: Recommendation two is about not collapsing all those human responses into a single label. Instead of taking the majority vote, you should keep the full distribution. So if sixty percent of people say “yes” and forty percent say “no,” that’s the target, not just “yes.”

Tom: And that’s a big shift. It means evaluating a model on how well it matches that distribution, not just whether it picks the most popular answer.

Lu: Right. And that connects to recommendation three: measure gradedness and uncertainty. Not just what people think, but how strongly they think it, and how confident they are.

Jane: They actually show that people often give intermediate slider values with high confidence. So it’s not that they’re unsure; it’s that the true answer is genuinely somewhere in the middle.

Meng: Okay, so we’re collecting richer data. But what about the tasks themselves? Are they even the right tasks?

Tom: That’s recommendation four. They say we need to ground our benchmarks in cognitive theory. Don’t just grab a classic psychology test like the Sally-Anne test and assume it measures everything about theory of mind.

Lu: Exactly. That test is designed for kids at a specific developmental stage. It’s not a comprehensive measure of adult social reasoning. So using it to claim an AI has theory of mind is a stretch.

Jane: And recommendation five is about making the tasks richer and more naturalistic. Real-world problems are messy. They involve ambiguity, incomplete information, and multiple cognitive processes at once.

Tom: So instead of simple multiple-choice questions, they want tasks that require integrating reasoning, perception, social understanding, maybe even planning. And they want to include stimuli that are deliberately ambiguous, because that’s where human cognition really shows its character.

Meng: So, if I’m building a benchmark, I need to collect lots of human data, keep the full distribution, measure graded confidence, ground it in theory, and make the tasks complex. That’s a lot of work.

Jane: It is, but it’s the only way to actually know if we’re building machines that think like us, rather than machines that are just really good at guessing what a benchmark designer thought was the right answer.

Tom: And that’s the hook for our final segment. Because if we do all this, what does it actually change in the real world?

Conclusion: Jane: So, Tom, we’ve covered the problems and the fixes. Let’s wrap up what “On Benchmarking Human-like Intelligence in Machines” really means for the future.

Tom: For me, the big takeaway is that we need to stop treating AI benchmarks like standardized tests with one right answer. They should be more like opinion polls, capturing the range of what people actually think.

Lu: And that has huge implications. If we build models that match human distributions, they’ll be better at collaborating with us. They’ll understand when we’re uncertain, when we’re joking, when we’re being sarcastic.

Meng: From a practical side, it means we need to invest in collecting high-quality human data. That’s expensive and time-consuming, but the paper argues it’s necessary if we want trustworthy AI.

Jane: And it’s not just about making AI more useful. It’s about understanding human intelligence better. By trying to match human behavior, we learn more about what that behavior actually is.

Tom: That’s the beautiful part. This paper is as much a call to action for cognitive science as it is for AI. It’s saying, “Let’s use these tools to understand ourselves.”

Lu: And the authors acknowledge it’s not easy. They talk about the challenges of scaling human data collection, and they admit that not every AI needs to be human-like. If you’re predicting protein structures, you don’t want human error.

Meng: Right, but for the tasks that matter for human interaction—moral judgment, social reasoning, humor—this is the way forward.

Jane: So we’re saying goodbye to this paper, but we’re taking its message with us. Next time someone claims their AI is human-like, we’re going to ask, “Compared to what?”

Tom: And with that, we’ll close out this discussion. Thanks for joining us on the channel. We’ll be back soon with another paper to pick apart.

Jane: Until then, keep questioning the benchmarks.

More episodes

← Home