OneMillion-Bench: How Far are Language Agents from Human Experts?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "OneMillion-Bench: How Far are Language Agents from Human Experts?".
Jane: The first text (A) is a technical description of a new benchmark, OneMillion-Bench (1M-Bench), designed to evaluate the capabilities of large language models (LLMs) as autonomous agents in complex,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're diving into this paper today, OneMillion-Bench: How Far are Language Agents from Human Experts? It comes from researchers like Qianyu Yang and Yang Liu, and the main idea is that current AI benchmarks just aren't tough enough to see if these models can actually do the kind of work people get paid for in real jobs.
Jane: Exactly. They’re saying we need to move past simple prompts and tests, because real professional tasks require agents to handle things like finding reliable sources and making decisions when evidence conflicts with itself. It matters because it asks if these models can actually create value in the economy right now, or if they are just good at answering questions for a test.
Lu: What's really interesting about this paper is how they set up these tasks across five huge areas: Law, Finance, Industry, Healthcare, and Natural Science. They aren't just asking for facts anymore; they need the AI to actually use external tools and figure out complex rules within those fields.
Meng: I’m curious about the actual testing part. How do they make sure these tasks are realistic enough for experts? Is it just having people write them down, or is there a specific way they validate that the task demands real professional depth?
Tom: It's not just one person writing them, Meng. The paper says they gathered time estimates from two to three senior experts in each domain to create these tasks. That way, it minimizes bias because you’re getting a mean estimation for how long these tasks actually take to complete, according to the authors > Expert Cost.
Jane: And they anchor that time estimate by using wage data that considers region and sector differences. They even adapted their approach for different countries, like using the U.S. Bureau of Labor Statistics data or tailoring things for Chinese tier-one cities > Wage Anchoring <ref:2603.07980#pg1>.
Lu: That wage anchoring is a big part of it because it’s about making sure the economic value they assign to these tasks is grounded in reality, not just some abstract number > Expert Cost, Wage Anchoring. It shows they are trying to quantify the real labor cost involved in these agentic capabilities.
Paper summary: Meng: So if an agent can handle one of these tasks—say, auditing a life insurance reserve valuation under IFRS seventeen—the authors calculate its economic value based on how long a senior actuary would take and what that person earns per hour > Expert Cost <ref:2603.07980#pg2,reserve valuation under IFRS 17>. That’s the core mechanism for measuring performance.
Tom: Right, so it's not just about whether the AI gets the right answer; it's about showing that the agent can perform a complex, multi-step process like an analyst constructing a high-pressure valuation model > Reference <ref:2603.07980#pg2>. It’s about sustained reasoning under strict professional constraints.
Jane: And they use this rubric system to make sure it’s not just about getting the final number right, but also if the whole reasoning process was sound and if the final solution is actually something someone could use in practice > rubric-based evaluation.
Lu: They are specifically looking for things like authoritative source retrieval, resolving conflicts between different pieces of evidence, applying domain-specific rules like legal precedents, and making those tough constraint decisions > Reference <ref:2603.07980#pg1>. That’s what they define as the necessary behaviors for an agent to be considered competent in this context.
Meng: I see the focus there. They are not just testing if an AI can talk; they are testing its ability to act like a specialist in a high-stakes situation where mistakes have real consequences > Reference <ref:2603.07980#pg2>. That shifts the whole goal of evaluation from surface-level correctness to actual reliability.
Tom: And when you look at the results, they found that even with search capabilities added, things still weren't perfect. They pointed out a gap between the Expert Scores and the Pass Rates > Reference <ref:2603.07980#pg3>. This means many models are only partially satisfying the requirements of a task rather than passing it completely, which is a real concern for reliability > Expert Cost, Wage Anchoring.
Jane: That gap tells us that simply adding search capability isn't enough on its own to make an agent reliable enough for professional work > Reference <ref:2603.07980#pg2>. They still need better mechanisms for how agents identify and use evidence correctly before they can truly perform well across multiple rubrics.
Lu: The paper shows that there is a clear stratification in performance, where the stronger models actually improve across different testing categories when search tools are available > Reference <ref:2603.07980#pg3>. It suggests that specialized systems or better reasoning architectures make a bigger difference than just throwing more search at a base model.
Paper summary: Meng: From an engineering standpoint, this confirms that we can’t just scale up the language size and expect the same level of professional competence without actually integrating robust research and retrieval mechanisms > Reference <ref:2603.07980#pg1>. The practical impact is that we need to build systems that are better at finding and trusting information, not just generating text.
Tom: So what’s the big picture for us? The title itself asks how far agents are from human experts, and this benchmark is trying to map out that distance using real economic stakes > Reference <ref:2603.07980#pg1>. It forces us to define what "doing the work necessary to answer well" actually looks like in a quantifiable way.
Jane: It’s about moving beyond the idea of an AI as a helpful assistant and starting to treat it as something that can reliably handle complex, multi-step, economically consequential decisions > Reference <ref:2603.07980#pg2>. The implication is that if we want agents to be used in high-stakes professional environments, we have to build them with this kind of rigorous evaluation in mind.
Lu: I think this benchmark really pushes the conversation toward building agents that are not just fluent but genuinely capable of performing the specific labor required by experts across those five fields > Reference <ref:2603.07980#pg1>. It’s about capability, not just linguistic skill.
Meng: It makes me think about how we train them next, because if they are failing on constraint decisions or evidence resolution, the training data or the fine-tuning needs to focus specifically on those failure points > Reference <ref:2603.07980#pg1>. We need targeted instruction for professional judgment.
Tom: So to wrap up this part of the paper, OneMillion-Bench: How Far are Language Agents from Human Experts? is showing that we need a new way to measure agentic maturity by focusing on grounded, compliant, and economically consequential decision-making > Reference <ref:2603.07980#pg1>. It’s about moving past simple correctness toward real-world performance in professional labor.
Jane: And it suggests that the next step for development isn't just making the models bigger, but making sure they actually understand how to use external information reliably and responsibly > Reference <ref:2603.07980#pg2>. That’s a big shift for how we think about agentic AI development.
Conclusion: Tom: So we’re wrapping up this look at OneMillion-Bench, and the title itself is pretty telling—"How Far are Language Agents from Human Experts?"
Jane: It really cuts right to the core question, doesn't it? They aren't just testing if an AI can answer a question; they’re measuring if it can actually perform a job that someone gets paid for.
Lu: The authors set up this massive test using real-world scenarios across five different fields, like law and finance, to see if the agent has the actual depth needed for that work.
Meng: It sounds like they’re trying to move past simple chat models into something that actually understands professional constraints and complex reasoning.
Lalam: From my side, this means our training needs to focus on those high-stakes decision points, because if an agent can't handle that level of complexity, it’s not ready for anything important.
Tom: Exactly. It shifts the goal from just being a good respondent to being something that can actually do the work necessary to answer well in a professional setting.
Jane: And they use this economic value metric—assigning real labor costs to the tasks—to show that this isn't just theoretical research; it’s about real-world consequences.
Lu: The way they structured the evaluation, focusing on evidence retrieval and resolving conflicts, tells us exactly what skills are missing in current models when they try to act like experts.
Meng: I think that gap between what a model *can* do and what an expert *has* to do is the main point here. It highlights where we need to focus our engineering efforts next.
Lalam: If we can get the AI consistently across all those rubrics, it means it’s capable of handling real professional demands reliably.
Tom: That reliability gap they found in their results suggests that just adding more search tools isn't enough on its own to make an agent truly competent for these tasks.
Jane: It really pushes us to think about building systems that prioritize grounded decisions and compliant behavior over just generating fluent text.
Lu: So, the paper is essentially saying we need a new way to measure maturity by focusing on economic consequences and professional reliability rather than just surface-level answers.
Meng: That means the next big hurdle for development isn't just making the models bigger, but making them fundamentally better at trusting and using external information correctly.
Humanlaya · BIGAI
cs.LG, cs.AI, cs.CL
Submitted: 2026-03-09
Updated: 2026-10-08
Code: https://github.com/MiniMax-AI/MiniMax-M2.1
Importance score: 89/100
The gist: The first text (A) is a technical description of a new benchmark, OneMillion-Bench (1M-Bench), designed to evaluate the capabilities of large language models (LLMs) as autonomous agents in complex,
Key concepts
- Agentic Evaluation
- This is the process of testing AI systems not just on single answers, but on their ability to perform multi-step tasks autonomously. It assesses if a model can use tools, retrieve external information, and resolve conflicts—the core skills needed for an agent to act like a human professional.
- Economic Grounding
- A unique feature where every task is assigned a real monetary value based on the time and cost of a senior professional doing it. This metric moves evaluation beyond simple accuracy to measure how valuable and practical an AI's output is in a business or legal context.
- Rubric-Based Protocol
- A detailed set of criteria used to score agent performance, focusing on four areas: factual accuracy, logical flow, practical feasibility, and professional compliance. This system ensures evaluation looks at the quality of reasoning and adherence to real-world standards rather than just surface-level correctness.
Terminology
Summary
The first text (A) is a technical description of a new benchmark, OneMillion-Bench (1M-Bench), designed to evaluate the capabilities of large language models (LLMs) as autonomous agents in complex, real-world professional scenarios. The second text (B) is an in-depth technical analysis explaining why Mean Squared Error (MSE) loss functions are fundamentally ill-suited for underwater image enhancement tasks due to perceptual and mathematical mismatches.
Since the request asks me to combine the summaries of both texts to better describe the paper, I must synthesize these two distinct concepts. However, it is crucial to note that Text A describes a benchmark for LLMs (Agentic Evaluation), while Text B describes a technical critique of an image processing loss function (MSE in Underwater Enhancement). These topics are entirely separate domains—one is Natural Language Processing/AI alignment, and the other is Computer Vision/Deep Learning optimization.
Given the instruction to describe the paper referring to OneMillion-Bench, I will focus on synthesizing Text A's content, as it directly pertains to the cited paper. I will incorporate elements from Text B only if they serve a meta-level purpose (e.g., drawing an analogy about suboptimal convergence
or perceptual disconnect,
although this is a stretch).
My primary focus will be on providing a long, detailed, and rigorous summary of the 1M-Bench paper based solely on Text A.
Detailed Research Synthesis: OneMillion-Bench (1M-Bench) Benchmark for Agentic Evaluation
The paper introduces OneMillion-Bench (1M-Bench) as a critical advancement in evaluating the capabilities of evolving language models, shifting the focus from simplistic, structured tasks to assessing agents' readiness for complex, economically consequential professional labor. The core premise is that existing benchmarks are insufficient because they fail to capture the demands of real-world agentic systems capable of multi-step reasoning and tool utilization.
I. Benchmark Design and Scope
1M-Bench is engineered to test agents across five high-stakes, domain-intensive sectors: Law, Finance, Industry, Healthcare, and Natural Science. The benchmark comprises 400 expert-curated tasks, each designed to simulate scenarios that require genuine professional depth.
Key Requirements for Task Complexity:
The tasks are not simple Q&A; they demand sophisticated agent behaviors including:
-
Authoritative Source Retrieval: Agents must be able to locate and utilize relevant, authoritative external information.
-
Evidence Resolution: They must navigate and resolve conflicts between disparate pieces of retrieved evidence.
-
Domain Rule Application: The system needs to correctly apply specific, domain-specific rules pertinent to the field (e.g., legal precedents, financial regulations).
-
Constraint Decision Making: Agents must make nuanced decisions under complex constraints.
II. Evaluation Protocol: Economic Grounding and Rubrics
A defining feature of 1M-Bench is its economic grounding. Every task is assigned a real-world monetary value, calculated based on the estimated time required for a senior professional to complete the task multiplied by their prevailing market hourly wage. Crucially, the total estimated economic value across all 400 tasks exceeds one million dollars, providing tangible quantification of agentic capability through labor cost metrics.
Evaluation is conducted using a rigorous rubric-based protocol designed specifically to mitigate reward hacking:
-
Factual Accuracy: Assessing correctness against established knowledge.
-
Logical Coherence: Evaluating the soundness and flow of the reasoning process.
-
Practical Feasibility: Determining if the proposed solution is actionable in a real-world context.
-
Professional Compliance: Verifying adherence to industry standards and ethical/legal constraints specific to the domain (e.g., compliance in Law or Finance).
This rubric-based approach ensures that evaluation moves beyond surface-level correctness toward assessing agentic reliability, professional depth, and practical readiness.
III. Agent Categories and Empirical Findings
1M-Bench systematically evaluates three distinct classes of language models:
-
Vanilla Models: Models operating without external tool integration.
-
Search Agents: Vanilla models augmented with web search capabilities.
-
Deep Research Agents: Specialized systems optimized for complex reasoning and long-context research tasks, often implying specialized architecture or fine-tuning for deep inference.
Key Empirical Insights from the Benchmark:
-
Top Performer Identification: The results indicate that CLAUDE-OPUS-4.6 emerges as the best overall performer among vanilla models and remains the top performer even when search capabilities are enabled.
-
The Efficacy of Search: Web Search is identified as an
Efficacy Amplifier.
While it significantly improves the Expert Score and Pass Rate for top models, it simultaneously introduces a risk: if agents lack robust mechanisms for evidence identification, search can introduce noisy or conflicting information. -
Agent Stratification: A clear stratification exists: strong models tend to improve across multiple rubrics when search is available, whereas weaker models often experience degradation in core reasoning and formatting skills.
-
The Pass Rate vs. Expert Score Gap: The research highlights a significant reliability gap between the two primary metrics. While many models achieve moderate Expert Scores (e.g., about45-50% on Global/CN tasks), their Pass Rates often remain substantially lower (frequently below about25%). This pattern suggests that model performance is frequently distributed as partially satisfying many rubrics rather than achieving a full competence threshold across a substantial portion of the test set.
-
Value of Specialization: Specialized Search Agents demonstrate drastically higher economic value compared to base models, reinforcing the necessity of integrating external tools for high-stakes tasks.
IV. Conclusion and Shift in Evaluation Paradigm
The overarching conclusion of 1M-Bench is a fundamental shift in how we measure agentic maturity. The benchmark demonstrates a significant reliability gap, showing that current models often fail to maintain the consistency and evidence-grounding necessary for autonomous professional labor. Therefore, 1M-Bench shifts the evaluative focus from mere surface-level correctness toward a framework that prioritizes grounded, compliant, and economically consequential decision-making as the true metric for agentic maturity. The ultimate goal is to propel language agents beyond simple answering toward systems expected not only to answer, but to do the work necessary to answer well.
Improvements for AI systems
-
Improve agentic reliability by shifting evaluation focus from aggregate accuracy to rubric compliance, as
Pass Rate complements Expert Score by exposing whether improvements reflect broad but shallow gains or genuinely push examples over the acceptance boundary (Expert Score(q) ≥ 0.7).
This ensures agents prioritize meeting domain-specific professional standards rather than superficial correctness. -
Enhance reasoning depth by implementing a multi-level aggregation strategy, allowing for analysis
by rubric type by averaging the normalized rubric scores of the same type across the dataset,
which revealswhere models generally achieve the highest scores on Structure and Formatting and Instructions Following, while Factual Information and Analytical Reasoning remain more challenging.
-
Improve evidence grounding by integrating web search as a conditional tool, recognizing that
Search amplifies underlying capabilities—models with better evidence filtering and planning can convert retrieval into score gains
but warns against its use in reasoning tasks where itmay introduce noisy or conflicting evidence.
-
Increase professional compliance by incorporating a negative rubrics scoring mechanism, which is designed to steer evaluation toward operational robustness by penalizing
violations of industry-specific norms or professional conduct, unsafe and harmful generation, factual hallucinations and lapses in expected foundational competencies such as instruction following.
-
Develop specialized performance tiers by differentiating model capabilities based on tool usage, as
Search agents deliver drastically higher economic value than that of the same base model,
establishing a clear trade-off between inference cost and high-value professional output.
Sources
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
- Measuring Massive Multitask Language Understanding
- Evaluation Framework for AI Systems in "the Wild"
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- GLM-5: from Vibe Coding to Agentic Engineering
- FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks