IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents
cs.CL, cs.AI, cs.HC
Submitted: 2026-01-10
Updated: 2026-09-23
Terminology
Sources
- Open Deep Search: Democratizing Search with Open-source Reasoning Agents
- DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
- Towards Asking Clarification Questions for Information Seeking on Task-Oriented Dialogues
- An Interactive Paradigm for Deep Research
- Retrieval-Augmented Generation for Large Language Models: A Survey
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- Qwen3 Technical Report
- Interaction-Driven Browsing: A Human-in-the-Loop Conceptual Framework Informed by Human Web Browsing for Browser-Using Agents
- Retrieval-Augmented Generation for AI-Generated Content: A Survey
- BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
- Trustworthiness in Retrieval-Augmented Generation Systems: A Survey
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering