DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports
cs.CL
Submitted: 2026-01-13
Updated: 2026-09-10
Code: https://github.com/imlrz/DeepResearch-Bench-II
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Deep Research Agents (DRA) aim to help users search the web, synthesize information, and deliver comprehensive investigative reports.
Terminology
Abstract
Deep Research Agents (DRA) aim to help users search the web, synthesize information, and deliver comprehensive investigative reports. Prior benchmarks often either under-evaluate a system's ability to produce meaningful insights and high-quality writing, or adopt coarse or LLM-defined criteria that are hard to verify and can diverge from human expert judgment. To address these issues, we introduce Deep Research Bench II, a new benchmark for evaluating DRAs. It contains 132 grounded research tasks across 22 domains; for each task, an agent must produce a research report that is evaluated by a set of 9,430 fine-grained binary rubrics in total, covering three dimensions: information recall, analysis, and presentation. All rubrics are derived from carefully selected expert-written investigative articles and are constructed through a four-stage LLM+human pipeline that combines automatic extraction with over 400 human-hours of expert review, ensuring that the criteria are verifiable and aligned with human expert judgment. We evaluate several state-of-the-art deep-research agents on Deep Research Bench II and find that even the strongest models satisfy fewer than 50% of the rubrics, revealing a substantial gap between current DRAs and human experts. We release the benchmark, evaluation scripts, and all rubrics at https://github.com/imlrz/DeepResearch-Bench-II to facilitate future research on deep-rearch agents.
Sources
- Language Models as Agent Models
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- Language Models are Few-Shot Learners
- Art or Artifice? Large Language Models and the False Promise of Creativity
- xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- Understanding DeepResearch via Reports
- Deep Research Bench: Evaluating AI Web Research Agents
- A Survey on LLM-as-a-Judge
- Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications
- Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
- Towards Long Context Hallucination Detection
- The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
- WebGPT: Browser-assisted question-answering with human feedback
- Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics
- Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
- MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- Evaluating LLMs with Multiple Problems at once
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering