Incremental Open-Ended Deep Research with Structured Harness
cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/modelscope/ms-agent2https:
Project page: https://ioedr-project.github.io
Terminology
Sources
- DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- Deep Researcher with Test-Time Diffusion
- Deep Research Agents: A Systematic Examination And Roadmap
- WebSailor: Navigating Super-human Reasoning for Web Agent
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- A Tale of Two Graphs: Separating Knowledge Exploration from Outline Structure for Open-Ended Deep Research
- Tongyi DeepResearch Technical Report
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- WebDancer: Towards Autonomous Information Seeking Agency
- Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models
- Deep Research: A Survey of Autonomous Research Agents
- AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool-Environment-Agent(TEA) Protocol
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering