F squared DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows
cs.CL
Submitted: 2026-09-17
Updated: 2026-09-19
License: http://creativecommons.org/licenses/by/4.0/
The gist: With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries.
Terminology
Abstract
With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, information retrieval, and answer generation. However, existing reward models (RMs) and evaluation benchmarks are primarily designed for static single-turn tasks, failing to capture the full-pipeline complexity of DeepSearch workflows. To address this limitation, we propose F2DR, a fine-grained full-pipeline DeepSearch reward framework. F2DR evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer, enabling comprehensive process-level assessment. We further construct DeepSearch RM-Bench, a dedicated benchmark for evaluating RMs in DeepSearch scenarios. Extensive experiments demonstrate that F2DR achieves significantly higher evaluation consistency than self-evaluation-based baselines, while DeepSearch RM-Bench exhibits strong discriminative capability across existing open-source RMs. We will publicly release the complete DeepSearch RM-Bench dataset soon.
Sources
- InternLM2 Technical Report
- IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
- RM-R1: Reward Modeling as Reasoning
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- Agentic Reinforced Policy Optimization
- Reward Reasoning Model
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
- RewardBench: Evaluating Reward Models for Language Modeling
- Inference-Time Scaling for Generalist Reward Modeling
- RewardBench 2: Advancing Reward Model Evaluation
- AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
- SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
- Tongyi DeepResearch Technical Report
- Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering