OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
cs.IR, cs.AI, cs.CL
Submitted: 2026-03-17
Updated: 2026-09-10
Code: https://github.com/TIGER-AI-Lab/OpenResearcher
License: http://creativecommons.org/licenses/by/4.0/
The gist: Training deep research agents requires long-horizon trajectories that interleave search, evidence aggregation, and multi-step reasoning.
Terminology
Abstract
Training deep research agents requires long-horizon trajectories that interleave search, evidence aggregation, and multi-step reasoning. However, existing data collection pipelines typically rely on proprietary web APIs, making large-scale trajectory synthesis costly, unstable, and difficult to reproduce. We present OpenResearcher, a reproducible pipeline that decouples one-time corpus bootstrapping from multi-turn trajectory synthesis and executes the search-and-browse loop entirely offline using three explicit browser primitives: search, open, and find, over a 15M-document corpus. Using GPT-OSS-120B as the teacher model, we synthesize over 97K trajectories, including a substantial long-horizon tail with 100+ tool calls. Supervised fine-tuning a 30B-A3B backbone on these trajectories achieves 54.8% accuracy on BrowseComp-Plus, a +34.0 point improvement over the base model, while remaining competitive on BrowseComp, GAIA, and xbench-DeepSearch. Because the environment is offline and fully instrumented, it also enables controlled analysis, where our study reveals practical insights into deep research pipeline design, including data filtering strategies, agent configuration choices, and how retrieval success relates to final answer accuracy. We release the pipeline, synthesized trajectories, model checkpoints, and the offline search environment at https://github.com/TIGER-AI-Lab/OpenResearcher.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- OpenThoughts: Data Recipes for Reasoning Models
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- WebSailor: Navigating Super-human Reasoning for Web Agent
- In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
- DeepSeek-V3 Technical Report
- WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
- AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
- APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
- Kimi K2: Open Agentic Intelligence
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG