HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam

arXiv:2602.13964 · cs.CL · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam".

Jane: The paper was written by Weiqi Zhai, Zhihai Wang, Jinghang Wang, Boyu Yang, Xiaogang Li et al. from Alibaba Group and Qwen Team, Alibaba Group.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the eye evaluation world. It's called "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam." Jane, I have to say, just the title alone tells you something big happened here.

Jane: It really does, Tom. So Humanity's Last Exam, or HLE, is this benchmark that everyone's been using to test the smartest eye models. It's got these incredibly hard questions across math, science, engineering, humanities. And the whole idea was to create something that wouldn't get saturated, something that would still be hard for the best models.

Tom: Right, and it's been cited everywhere. People use it to claim which model is smarter, which one reasons better. But then this paper comes along and says, hold on, maybe some of those questions are just wrong.

Jane: Exactly. And that's what HLE-Verified is about. The authors went through all two thousand five hundred questions in HLE and systematically checked them. They broke each question into three parts, the problem statement, the final answer, and the rationale, which is the explanation for how you get to the answer.

Tom: And what did they find?

Jane: Well, only six hundred sixty-eight questions were verified as completely correct as they were. That's about a quarter of the whole dataset. Another one thousand one hundred forty-three questions had errors but could be fixed. And six hundred eighty-nine questions were so uncertain that they couldn't be verified at all.

Tom: Wait, so more than half the questions needed fixing or couldn't be verified? That's a huge deal.

Jane: It is. And think about what that means for all the leaderboards people have been looking at. If a question has the wrong answer key, then a model that gives the right answer gets marked wrong. And a model that gives the wrong answer, matching the bad key, gets marked right.

Tom: So the rankings we've been seeing might be partly measuring how good models are at matching bad answers.

Jane: Precisely. And that's why this paper matters so much. It's not just about finding errors. It's about making the benchmark trustworthy again. The authors didn't just throw out the bad questions. They fixed the ones they could, and they documented everything so the community can keep improving it.

Tom: I love that they released all the metadata too. It's not just a new dataset, it's a whole system for auditing benchmarks. That feels like a shift in how we should think about evaluation.

Jane: Totally. And we're going to get into exactly how they did the verification and what happened when they tested models on the fixed version. That's where it gets really interesting.

Tom: Can't wait. Let's dig into the summary next.

Summary: Tom: So Jane, we've established that HLE-Verified is a big deal because it found a lot of problems in the original HLE benchmark. But what exactly did the authors do? What's the actual process here?

Jane: Great question. They built a two-stage pipeline. Stage one is verification. They took every single question and had domain experts look at the problem, the answer, and the rationale separately. Each component gets a label, valid, invalid, or uncertain.

Tom: And they didn't just rely on human experts, right?

Jane: Right. They also used multiple eye models to try solving each question independently. If several strong models all fail on a question, that's a red flag that something might be wrong with the question itself, not the models. But here's the key, the model results are just a signal. They're not the final say. Human experts make the final call.

Tom: So stage one gives you the gold subset, the six hundred sixty-eight questions that are fine as is. Then what happens to the rest?

Jane: Stage two is revision. For the questions that are flawed but fixable, they have two independent teams of experts propose corrections. Then they use eye models to check the stability of those corrections. And finally, an internal adjudicator picks the best version or synthesizes one.

Tom: And there's a really important constraint there, right? They can't just rewrite the question into something easier.

Jane: Exactly. The whole point is to preserve the original evaluation intent. If a question is testing whether a model can derive a specific formula, you can't change it into a different formula question. You fix the errors, you make the question well-posed, but you don't change what it's testing.

Tom: And what about the six hundred eighty-nine uncertain ones?

Jane: Those aren't discarded either. They're kept in a separate category with documentation about why they're uncertain and what kind of expertise would be needed to resolve them. So the dataset has three clear buckets, verified, revised, and uncertain. Everything is transparent.

Tom: That's really thoughtful. Instead of just deleting the hard cases, they're saying, here's what we couldn't resolve, and here's what you'd need to fix it.

Jane: And that's what makes this a methodological contribution, not just a dataset release. They've created a framework for auditing benchmarks that other people can apply to other benchmarks too.

Tom: So what happened when they actually tested models on the revised benchmark? I heard the numbers are pretty dramatic.

Jane: They are. On the items that were revised, models saw their accuracy jump by thirty to forty percentage points on average. That's not a small effect. That's a massive effect.

Tom: Thirty to forty points. That means a model that looked like it was getting twenty percent on those questions was actually getting fifty or sixty percent once the questions were fixed.

Jane: Exactly. And that tells us the original benchmark was systematically underestimating model capability. The errors in the benchmark were creating a ceiling that didn't really exist.

Tom: So the models were smarter than the benchmark was giving them credit for.

Jane: In many cases, yes. And we're going to talk about what that means for how we should interpret all those HLE leaderboards, and also what it means for model confidence. But first, let's look at the specific improvements they made.

Improvements: Tom: So we know the accuracy numbers went way up on the revised questions. But what kinds of errors were they actually fixing? What did they find when they dug into the questions?

Jane: They built a really detailed taxonomy of defects. Nineteen different categories. Five for problems, ten for rationales, and four for answers. And the patterns they found are fascinating.

Tom: Give me the highlights.

Jane: Well, the most common answer-level defect is just a plain incorrect answer. Across all subjects, that's the dominant problem. In biology and medicine, it's almost ninety-seven percent of the answer defects. So the question is fine, but the answer key is just wrong.

Tom: That's brutal. A model could give a perfect answer and get marked wrong because the key says something else.

Jane: Exactly. And then at the rationale level, the most common issues are missing information and format semantic errors. So the reasoning steps have gaps, or the notation is so messed up that it changes the meaning. And at the problem level, format semantic errors dominate in math, chemistry, and computer science.

Tom: So it's not that the questions are conceptually broken most of the time. It's that the representation is broken, or the answer key is wrong.

Jane: Right. And they have some great case studies that show this. There's one in computer science about speculative decoding. The question asks what the expected acceptance rate is when the same model is both draft and target. The original answer claimed it should be less than one because of GPU kernel differences.

Tom: But that's mixing up implementation details with theoretical properties.

Jane: Exactly. The correct answer is exactly one, because if the draft and target are the same model, they should agree perfectly. The GPU kernel stuff is just numerical noise. So the original answer was teaching the wrong lesson.

Tom: And they fixed that by separating the theory from the hardware artifacts.

Jane: Yes. Another example from chemistry had incorrect molecular mass constants, making the whole system mathematically inconsistent. And a physics example had a sign error that gave a negative Rydberg energy, which makes no physical sense.

Tom: These aren't subtle issues. These are fundamental errors in the benchmark itself.

Jane: And that's why the accuracy gains are so large. When you fix the answer key, models that were already giving the right answer suddenly get credit for it. The paper shows that on the revised subset, models like GPT-five point two jumped from about fourteen percent accuracy to over fifty-two percent.

Tom: That's a thirty-eight point jump just from fixing the benchmark.

Jane: Yes. And it's consistent across all the models they tested. Every single model improved. Some by thirty points, some by almost forty. And the calibration error also dropped significantly.

Tom: Calibration error, that's how well a model's confidence matches its actual correctness, right?

Jane: Exactly. And that's the next thing we should talk about, because they found something really interesting about model confidence and noisy questions.

Conclusion: Tom: So we've covered the verification process, the error taxonomy, and the big accuracy gains. But there's one more finding that I think is really important, and it's about model confidence.

Jane: Yes. They looked at how confident models are when they answer questions. And they found that on questions with errors in the problem statement, models tend to be less confident. When those questions are fixed, model confidence goes up.

Tom: So the models can sense something is off, even if they can't always identify what it is.

Jane: Exactly. On the full dataset, the confidence shift is near zero because most questions are unchanged. But on the subset with problem-level errors, confidence increases by anywhere from two to eleven points after repair.

Tom: That's a really useful diagnostic. If a model is consistently less confident on a question, that might be a signal that the question itself is flawed.

Jane: And that's actually one of the most practical takeaways from this paper. Confidence could be used as a screening tool for finding noisy items in other benchmarks too.

Tom: So let's wrap this up. The paper is "HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam." And the big picture is that benchmarks aren't sacred. They need maintenance. They need auditing. And this paper provides a blueprint for how to do that.

Jane: And the impact goes beyond just HLE. Any benchmark that's used to compare models could benefit from this kind of component-wise verification. MMLU, GPQA, SWE-bench, they all have the same potential issues.

Tom: And the fact that they released the uncertain set with documentation means the community can keep working on it. It's not a one-time fix. It's an ongoing process.

Jane: Right. And I think that's the real message here. Evaluation is not a solved problem. It's something we have to keep working on, especially as models get better and benchmarks get harder.

Tom: Well said, Jane. That's a great note to end on. Thanks to everyone for listening. We'll be back next time with another paper. Until then, keep questioning the benchmarks.

Jane: And keep questioning the questions. See you all next time.

Weiqi Zhai, Zhihai Wang, Jinghang Wang, Boyu Yang, Xiaogang Li, Xander Xu, Bohan Wang, Peng Wang, Xingzhe Wu, Anfeng Li, Qiyuan Feng, Yuhao Zhou, Taolin Han, Wenjie Luo, Yiyuan Li, Xiang Zheng, Yaxuan Wang, Ruixiang Luo, Guojie Lin, Peiyao Xiao, Chengliang Xu, Ben Wang, Zeyu Wang, Zichao Chen, Jianan Ye, Yijie Hu, Jialong Chen, Zongwen Shen, Yuliang Xu, An Yang, Bowen Yu, Dayiheng Liu, Junyang Lin, Hu Wei, Que Shen, Bing Zhao

Alibaba Group · Qwen Team, Alibaba Group

cs.CL

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 14 pages, 10 figures

Code: https://github.com/lhl/hle-gpqa-error-claims

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: "To address this challenge, we introduce HLE-Verified, a verified and revised version of HLE accompanied by a transparent, component-wise verification protocol and fine-grained error taxonomy." The

Key concepts

Humanity's Last Exam (HLE)
HLE is a difficult benchmark used to test AI models. It contains challenging questions across various subjects like math, science, and humanities. Its purpose was to create a standard that would not become easily saturated by advanced AI models.
HLE-Verified
This refers to the paper and methodology used to systematically audit HLE. The authors reviewed all 2500 questions, identifying errors in the answer keys or flaws in the problem statements, creating a framework for ensuring benchmark trustworthiness.
Verification and Revision Pipeline
This is a two-stage process. First, domain experts label components of questions (problem, answer, rationale) as valid or invalid. Second, fixable errors are corrected by expert teams while strictly preserving the original evaluation intent.
Model Confidence
This measures how well a model's certainty aligns with its actual correctness. The study found that models tend to be less confident on questions containing errors in the problem statement, and this confidence increases after repair.

Terminology

Summary

Summary

The paper introduces HLE-Verified, a verified and revised version of the Humanity’s Last Exam (HLE) benchmark, constructed via a transparent, component-wise verification protocol and a fine-grained error taxonomy. The motivation is that HLE, while widely used for evaluating frontier language models, contains a non-trivial number of noisy items (e.g., ambiguous statements, incorrect answers, or mismatched rationales) that can systematically bias evaluation results and distort cross-model comparisons. The authors state: To address this challenge, we introduce HLE-Verified, a verified and revised version of HLE accompanied by a transparent, component-wise verification protocol and fine-grained error taxonomy.

The construction follows a two-stage validation-and-repair workflow resulting in a unified certified benchmark. In Stage I, each item is subjected to binary validation on the problem and final answer dimensions (with rationale used as an auxiliary consistency signal), combining domain-expert review and model-based cross-checks. This stage yields 668 items verified as correct. In Stage II, items identified as flawed but fixable are systematically revised under strict constraints that preserve the original evaluation intent, through dual independent expert repairs, model-assisted consistency auditing, and final expert adjudication, resulting in 1,143 revised-and-certified items. The remaining 689 items are released as a documented uncertain set with explicit uncertainty sources and required expertise tags for future community refinement.

The paper's contributions are: (1) releasing HLE-Verified comprising the three subsets above; (2) introducing a systematic two-stage verification-and-revision framework for post-release benchmark auditing; and (3) benchmarking eight state-of-the-art LLMs on HLE vs. HLE-Verified, demonstrating that verification materially alters measured performance.

Key methodological details: Each item is decomposed into three annotatable components: the problem (statement plus image, if present), the final answer, and the rationale (reference solution). The problem and final answer are primary objects of correctness assessment, while the rationale serves as diagnostic support. Stage I integrates three sources of evidence: external domain-expert screening, model-assisted replication checks (pass@8), and internal expert adjudication. Items enter the gold subset only when both problem and answer are judged unproblematic and no high-risk ambiguity is identified. Stage II focuses on domains where correctness can be reliably adjudicated (mathematics, physics, chemistry, biomedicine, and computer science), generating independent repair proposals followed by expert convergence, with revisions performed under the constraint that the original evaluation objective and reasoning target are preserved.

The paper defines a 19-category defect taxonomy: 5 problem-level errors (Q1 Semantic Error, Q2 Knowledge Error, Q3 Missing Information, Q4 Theoretical Invalidity, Q5 Format Semantic Error), 10 rationale-level errors (S1 Non-Redundancy Violation, S2 Circular Reasoning, S3 Empirical Soundness Violation, S4 Step Inconsistency, S5 Domain Misapplication, S6 Overconfidence Bias, S7 Missing Prerequisite, S8 Deceptive Similarity, S9 Multi-Solution Inconsistency, S10 Format Semantic Error), and 4 answer-level errors (A1 Incorrect Answer, A2 Incomplete Answer, A3 Ambiguous/Ill-defined Answer, A4 Format Semantic Error).

Statistical analysis findings: Across the problematic subset, the Problem field exhibits the highest reliability, while reliability decreases at the Answer level (only about half of answers are fully valid), and the most pronounced degradation appears in the Rationale component (invalid rationales outnumber valid ones). The uncertain label constitutes a non-trivial fraction across all components. Cross-domain variation shows that Math and Biology/Medicine exceed 92% problem validity, while Physics shows only 30.0% valid problem statements. At the Answer level, Computer Science/AI (65.1%) and Chemistry (67.3%) achieve the highest validity rates. For the Rationale component, invalid rationales dominate in Math (65.9%), Biology/Medicine (66.4%), and Chemistry (70.3%). The dominant defect types are: Incorrect Answer dominates answer-level defects across all subjects (ranging from 69.4% to 97.2%); rationale defects are primarily structural incompleteness or format-induced semantic errors; and problem-level defects are dominated by format semantic errors in Mathematics (40.0%), Chemistry (56.5%), and Computer Science (94.6%), while Physics is led by Semantic Error (57.1%) and Biology/Medicine by Knowledge Error (33.3%).

Experimental results: The paper evaluates eight frontier models: GPT-5.2-Thinking, Gemini3-Pro-Preview, Claude-Opus4.5, Claude-Opus4.6, Grok-4.1 (fast-reasoning), DeepSeek-V3.2-Thinking, Qwen3-Max-Thinking, and Qwen3.5-Plus. All models were evaluated using the same system prompt recommended by HLE official guidelines, with five independent rollouts per item, reporting avg5 accuracy and calibration error. On the Revised Subset (items where at least one of problem/answer fields was revised), accuracy increases by: Gemini-3-pro (+29.94), GPT-5.2 (+38.04), Claude-Opus4.5 (+32.94), Grok-4.1 fast-reasoning (+34.82), Claude-Opus4.6 (+30.13), and DeepSeek-V3.2 (+39.58) percentage points. Calibration error decreases consistently after revision (e.g., GPT-5.2: 63→28; DeepSeek-V3.2: 70→28; Grok-4.1: 83→47). On the Full Set, accuracy increases by: Gemini-3-pro (+7.58), GPT-5.2 (+9.95), Claude-Opus4.5 (+8.68), Grok-4.1 fast-reasoning (+9.26), Claude-Opus4.6 (+7.75), DeepSeek-V3.2 (+10.79), and Qwen3-Max-Thinking (+8.92) percentage points, with calibration error also decreasing across models. The paper states: "The improvement is particularly pronounced on items where the original HLE problem statement and/or reference answer is erroneous: on this subset, the models achieve an average accuracy increase of 30–40 percentage points."

Confidence analysis: The paper studies the relationship between model confidence and item quality by comparing confidence statistics before and after verification on items whose problem statements are flagged as erroneous. The mean confidence shift ∆Conf = E[cVerified − cRaw] increases consistently after repair on the Problem-Error Subset across all evaluated models, with gains ranging from roughly +1.83 to +11.08 confidence points (absolute), while Full-Set shifts are near zero. The authors conclude: "These results suggest that statement-level noise in raw HLE may not only depress accuracy but also reduce model confidence in a systematic way. As a result, confidence (and related calibration signals) could be a useful diagnostic for flagging potentially noisy or ambiguous items."

Case studies illustrate structural defect patterns: (1) Theoretical–Implementation Confusion in Computer Science, where a speculative decoding sanity-check conflated implementation-level numerical variability with a theoretical property, with the correct answer being exactly 1 under identical models; (2) Numerical Inconsistency in Stoichiometric Constraint in Chemistry, where incorrect molecular mass constants made the system unsatisfiable, corrected to restore internal coherence; (3) Sign Inconsistency in Exciton Energy in Physics, where confusion among bandgap energy, resonance peak, and binding energy introduced a negative Rydberg energy, corrected to restore physical coherence.

The paper concludes: "We introduced HLE-Verified, a verified and revised benchmark intended to strengthen the scientific reliability of HLE-based evaluations. Our work makes benchmark flaws measurable, provides a transparent correction pipeline, and quantifies how such flaws bias reported accuracy and calibration. Beyond this release, the disputed set provides a roadmap for community-driven improvements. We expect that a continuously maintained verification process, with structured metadata and clear contribution guidelines, can make exam-style benchmarks more robust and more informative for tracking real progress in language model reasoning." Data is available at: https://huggingface.co/datasets/skylenage/HLE-Verified.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems:

Improvement: Add a pre-evaluation filter that detects and flags potentially flawed benchmark items before scoring.

Implementation:

  • Integrate a component-wise validity checker (Problem/Answer/Rationale) using the paper's 19-category defect taxonomy

  • Before final scoring, run items through a lightweight classifier trained on the HLE-Verified annotations to identify likely-erroneous items

  • Report both raw accuracy and verified accuracy (excluding flagged items) to provide a more reliable capability estimate

Resulting capability: The system can now distinguish genuine model failures from benchmark artifacts, reducing false-negative evaluations by 30-40% on noisy items.


Summary of what the improved system can do: It can now evaluate models more accurately by filtering benchmark noise, produce more reliable answers on ambiguous inputs, self-correct for common error patterns, provide domain-aware confidence, and validate new benchmarks—all leading to more trustworthy AI performance measurement and deployment.

Abstract

Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses have raised concerns that HLE contains a non-trivial number of noisy items, which can bias evaluation results and distort cross-model comparisons. To address this challenge, we introduce HLE-Verified, a verified and revised version of HLE with a transparent verification protocol and fine-grained error taxonomy. Our construction follows a two-stage validation-and-repair workflow resulting in a certified benchmark. In Stage I, each item undergoes binary validation of the problem and final answer through domain-expert review and model-based cross-checks, yielding 668 verified items. In Stage II, flawed but fixable items are revised under strict constraints preserving the original evaluation intent, through dual independent expert repairs, model-assisted auditing, and final adjudication, resulting in 1,143 revised-and-certified items. The remaining 689 items are released as a documented uncertain set with explicit uncertainty sources and expertise tags for future refinement. We evaluate eight state-of-the-art language models on HLE and HLE-Verified, observing an average absolute accuracy gain of 7--10 percentage points on HLE-Verified. The improvement is particularly pronounced on items where the original problem statement and/or reference answer is erroneous, with gains of 30--40 percentage points. Our analyses further reveal a strong association between model confidence and the presence of errors in the problem statement or reference answer, supporting the effectiveness of our revisions. Overall, HLE-Verified improves HLE-style evaluations by reducing annotation noise and enabling more faithful measurement of model capabilities. Data is available at: https://huggingface.co/datasets/skylenage/HLE-Verified

Sources

Related papers