OneMillion-Bench: How Far are Language Agents from Human Experts?

summary

Video file (mp4)

The gist

The first text (A) is a technical description of a new benchmark, OneMillion-Bench (1M-Bench), designed to evaluate the capabilities of large language models (LLMs) as autonomous agents in complex,

In short

OneMillion-Bench (1M-Bench) was created to rigorously test if large language models can function as autonomous agents capable of complex, real-world professional work. By simulating 400 high-stakes tasks across Law, Finance, Healthcare, etc., the benchmark measures agent reliability through economic grounding and strict rubrics. The findings reveal a significant gap between models that merely answer questions and those that can reliably perform grounded, compliant decision-making required for expert labor.

Key concepts

Agentic Evaluation
This is the process of testing AI systems not just on single answers, but on their ability to perform multi-step tasks autonomously. It assesses if a model can use tools, retrieve external information, and resolve conflicts—the core skills needed for an agent to act like a human professional.
Economic Grounding
A unique feature where every task is assigned a real monetary value based on the time and cost of a senior professional doing it. This metric moves evaluation beyond simple accuracy to measure how valuable and practical an AI's output is in a business or legal context.
Rubric-Based Protocol
A detailed set of criteria used to score agent performance, focusing on four areas: factual accuracy, logical flow, practical feasibility, and professional compliance. This system ensures evaluation looks at the quality of reasoning and adherence to real-world standards rather than just surface-level correctness.

Terminology used across episodes

This episode discusses

The paper

$OneMillion-Bench: How Far are Language Agents from Human Experts? · Read on arXiv

Humanlaya · BIGAI

As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce OneMillion-Bench (OMB), a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focusing on expert-level problems to ensure meaningful differentiation across agents. Together, OMB provides a unified testbed for assessing agentic reliability, professional depth, and an indicator of practical readiness in domain-intensive scenarios.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "OneMillion-Bench: How Far are Language Agents from Human Experts?".

Jane: The first text (A) is a technical description of a new benchmark, OneMillion-Bench (1M-Bench), designed to evaluate the capabilities of large language models (LLMs) as autonomous agents in complex,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're diving into this paper today, OneMillion-Bench: How Far are Language Agents from Human Experts? It comes from researchers like Qianyu Yang and Yang Liu, and the main idea is that current AI benchmarks just aren't tough enough to see if these models can actually do the kind of work people get paid for in real jobs.

Jane: Exactly. They’re saying we need to move past simple prompts and tests, because real professional tasks require agents to handle things like finding reliable sources and making decisions when evidence conflicts with itself. It matters because it asks if these models can actually create value in the economy right now, or if they are just good at answering questions for a test.

Lu: What's really interesting about this paper is how they set up these tasks across five huge areas: Law, Finance, Industry, Healthcare, and Natural Science. They aren't just asking for facts anymore; they need the AI to actually use external tools and figure out complex rules within those fields.

Meng: I’m curious about the actual testing part. How do they make sure these tasks are realistic enough for experts? Is it just having people write them down, or is there a specific way they validate that the task demands real professional depth?

Tom: It's not just one person writing them, Meng. The paper says they gathered time estimates from two to three senior experts in each domain to create these tasks. That way, it minimizes bias because you’re getting a mean estimation for how long these tasks actually take to complete, according to the authors > Expert Cost.

Jane: And they anchor that time estimate by using wage data that considers region and sector differences. They even adapted their approach for different countries, like using the U.S. Bureau of Labor Statistics data or tailoring things for Chinese tier-one cities > Wage Anchoring <ref:2603.07980#pg1>.

Lu: That wage anchoring is a big part of it because it’s about making sure the economic value they assign to these tasks is grounded in reality, not just some abstract number > Expert Cost, Wage Anchoring. It shows they are trying to quantify the real labor cost involved in these agentic capabilities.

Paper summary: Meng: So if an agent can handle one of these tasks—say, auditing a life insurance reserve valuation under IFRS seventeen—the authors calculate its economic value based on how long a senior actuary would take and what that person earns per hour > Expert Cost <ref:2603.07980#pg2,reserve valuation under IFRS 17>. That’s the core mechanism for measuring performance.

Tom: Right, so it's not just about whether the AI gets the right answer; it's about showing that the agent can perform a complex, multi-step process like an analyst constructing a high-pressure valuation model > Reference <ref:2603.07980#pg2>. It’s about sustained reasoning under strict professional constraints.

Jane: And they use this rubric system to make sure it’s not just about getting the final number right, but also if the whole reasoning process was sound and if the final solution is actually something someone could use in practice > rubric-based evaluation.

Lu: They are specifically looking for things like authoritative source retrieval, resolving conflicts between different pieces of evidence, applying domain-specific rules like legal precedents, and making those tough constraint decisions > Reference <ref:2603.07980#pg1>. That’s what they define as the necessary behaviors for an agent to be considered competent in this context.

Meng: I see the focus there. They are not just testing if an AI can talk; they are testing its ability to act like a specialist in a high-stakes situation where mistakes have real consequences > Reference <ref:2603.07980#pg2>. That shifts the whole goal of evaluation from surface-level correctness to actual reliability.

Tom: And when you look at the results, they found that even with search capabilities added, things still weren't perfect. They pointed out a gap between the Expert Scores and the Pass Rates > Reference <ref:2603.07980#pg3>. This means many models are only partially satisfying the requirements of a task rather than passing it completely, which is a real concern for reliability > Expert Cost, Wage Anchoring.

Jane: That gap tells us that simply adding search capability isn't enough on its own to make an agent reliable enough for professional work > Reference <ref:2603.07980#pg2>. They still need better mechanisms for how agents identify and use evidence correctly before they can truly perform well across multiple rubrics.

Lu: The paper shows that there is a clear stratification in performance, where the stronger models actually improve across different testing categories when search tools are available > Reference <ref:2603.07980#pg3>. It suggests that specialized systems or better reasoning architectures make a bigger difference than just throwing more search at a base model.

Paper summary: Meng: From an engineering standpoint, this confirms that we can’t just scale up the language size and expect the same level of professional competence without actually integrating robust research and retrieval mechanisms > Reference <ref:2603.07980#pg1>. The practical impact is that we need to build systems that are better at finding and trusting information, not just generating text.

Tom: So what’s the big picture for us? The title itself asks how far agents are from human experts, and this benchmark is trying to map out that distance using real economic stakes > Reference <ref:2603.07980#pg1>. It forces us to define what "doing the work necessary to answer well" actually looks like in a quantifiable way.

Jane: It’s about moving beyond the idea of an AI as a helpful assistant and starting to treat it as something that can reliably handle complex, multi-step, economically consequential decisions > Reference <ref:2603.07980#pg2>. The implication is that if we want agents to be used in high-stakes professional environments, we have to build them with this kind of rigorous evaluation in mind.

Lu: I think this benchmark really pushes the conversation toward building agents that are not just fluent but genuinely capable of performing the specific labor required by experts across those five fields > Reference <ref:2603.07980#pg1>. It’s about capability, not just linguistic skill.

Meng: It makes me think about how we train them next, because if they are failing on constraint decisions or evidence resolution, the training data or the fine-tuning needs to focus specifically on those failure points > Reference <ref:2603.07980#pg1>. We need targeted instruction for professional judgment.

Tom: So to wrap up this part of the paper, OneMillion-Bench: How Far are Language Agents from Human Experts? is showing that we need a new way to measure agentic maturity by focusing on grounded, compliant, and economically consequential decision-making > Reference <ref:2603.07980#pg1>. It’s about moving past simple correctness toward real-world performance in professional labor.

Jane: And it suggests that the next step for development isn't just making the models bigger, but making sure they actually understand how to use external information reliably and responsibly > Reference <ref:2603.07980#pg2>. That’s a big shift for how we think about agentic AI development.

Conclusion: Tom: So we’re wrapping up this look at OneMillion-Bench, and the title itself is pretty telling—"How Far are Language Agents from Human Experts?"

Jane: It really cuts right to the core question, doesn't it? They aren't just testing if an AI can answer a question; they’re measuring if it can actually perform a job that someone gets paid for.

Lu: The authors set up this massive test using real-world scenarios across five different fields, like law and finance, to see if the agent has the actual depth needed for that work.

Meng: It sounds like they’re trying to move past simple chat models into something that actually understands professional constraints and complex reasoning.

Lalam: From my side, this means our training needs to focus on those high-stakes decision points, because if an agent can't handle that level of complexity, it’s not ready for anything important.

Tom: Exactly. It shifts the goal from just being a good respondent to being something that can actually do the work necessary to answer well in a professional setting.

Jane: And they use this economic value metric—assigning real labor costs to the tasks—to show that this isn't just theoretical research; it’s about real-world consequences.

Lu: The way they structured the evaluation, focusing on evidence retrieval and resolving conflicts, tells us exactly what skills are missing in current models when they try to act like experts.

Meng: I think that gap between what a model *can* do and what an expert *has* to do is the main point here. It highlights where we need to focus our engineering efforts next.

Lalam: If we can get the AI consistently across all those rubrics, it means it’s capable of handling real professional demands reliably.

Tom: That reliability gap they found in their results suggests that just adding more search tools isn't enough on its own to make an agent truly competent for these tasks.

Jane: It really pushes us to think about building systems that prioritize grounded decisions and compliant behavior over just generating fluent text.

Lu: So, the paper is essentially saying we need a new way to measure maturity by focusing on economic consequences and professional reliability rather than just surface-level answers.

Meng: That means the next big hurdle for development isn't just making the models bigger, but making them fundamentally better at trusting and using external information correctly.

More episodes

← Home