OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: 9 pages, 4 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Search-augmented reasoning remains difficult for small language models.
Terminology
Abstract
Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from trained teachers offers a promising direction, but suffers from two issues: (1) high-quality multi-turn search trajectories depend on dynamic retriever responses, making SFT data prohibitively expensive to collect at scale; (2) task-specifically trained teachers incur substantial training cost, while directly applying OPD with an off-the-shelf teacher without task-specific fine-tuning constrains the student to the teacher's performance ceiling and suffers from severe training instability. We propose OPDSearch+, the first distillation paradigm that requires no teacher fine-tuning for search-augmented reasoning. We investigate the role of a frozen off-the-shelf instruct model as the teacher in on-policy distillation, and reveal a key insight: the teacher reshapes the student's policy distribution so that subsequent RL converges to a superior solution that RL alone cannot reach. In stage one, the student interacts with a live search engine and is distilled via a per-position forward KL objective, transferring reasoning decomposition and evidence integration skills without any task-specific teacher training. In stage two, RL refines the distilled student from a richer behavioral foundation, achieving performance that RL alone cannot reach from scratch. Across seven QA benchmarks, OPDSearch+ with a 3B model consistently outperforms all prior 3B RL baselines, achieving gains of 13.1% on HotpotQA and 8.5% on 2WikiMultihopQA.
Sources
- Group-in-Group Policy Optimization for LLM Agent Training
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- Distilling the Knowledge in a Neural Network
- Healthcare AI GYM for Medical Agents
- Entropy-Aware On-Policy Distillation of Language Models
- Distilling LLM Agent into Small Models with Retrieval and Code Tools
- Surgical Post-Training: Proximal On-Policy Distillation for Reasoning with Knowledge Retention
- Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence
- SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning
- WebGPT: Browser-assisted question-answering with human feedback
- Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards
- HybridFlow: A Flexible and Efficient RLHF Framework
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
- KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
- Reinforcement-aware Knowledge Distillation for LLM Reasoning
- Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection