Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
Haoze Wu, Chuqiao Kuang, Tianyi Zhuang, Xiaoguang Li
The Hong Kong University of Science and Technology · Huawei Technologies Ltd.
cs.LG, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Work in progress
Code: https://github.com/hkust-nlp/SSPO
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: The paper "Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents" addresses the challenge of sparse reward supervision in training deep search agents, which
Terminology
Summary
The paper Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
addresses the challenge of sparse reward supervision in training deep search agents, which operate over trajectories spanning dozens of steps but receive only a single binary outcome reward. The authors identify that standard reinforcement learning (RL) provides insufficient credit assignment in this setting.
The paper proposes two main contributions to resolve the tension between on-policy self-distillation (OPSD) and the information asymmetry inherent in multi-turn search agents. First, they construct Evidence Anchors, which are concise, step-level evidence snippets extracted from the web that capture information needed to answer a question without revealing the answer path.
These serve as privileged information for a self-teacher. Second, they propose Step-Level Self-Distilled Policy Optimization (SSPO), which converts teacher–student disagreement into step-level advantage weights within GRPO, applied exclusively to incorrect trajectories.
This design decouples the update direction (determined by outcome reward) from the update magnitude (modulated by the teacher at each step), leaving correct trajectories untouched to preserve diversity.
The key findings and results are as follows:
-
Problem with direct distillation: The paper shows that directly matching the teacher distribution collapses the student's tool use (from 17.7 to 3.5 average turns) and underperforms even GRPO, confirming that decoupling update magnitude from direction is essential.
-
Superior performance: On Qwen3-8B, SSPO consistently outperforms GRPO across BrowseComp, GAIA, and FRAMES benchmarks. Notably,
SSPO trained for 100 steps already surpasses GRPO trained for 200 steps,
highlighting superior sample efficiency. -
Quantitative gains: Using the average score across three benchmarks, GRPO achieves a +2.4 improvement over the cold-start baseline, whereas SSPO delivers a markedly larger gain of +4.8. Specifically, on BrowseComp, SSPO achieves 15.7 accuracy vs. GRPO's 13.6; on GAIA, 49.3 vs. 47.3; and on FRAMES, 73.0 vs. 69.8.
-
Minimal overhead: The additional teacher forward pass accounts for only about 5% of the total step time, making the overhead negligible relative to performance gains.
-
Granularity matters: The ablation study shows that step-level advantage weights consistently outperform token-level counterparts, whose fine-grained signals are misaligned with the natural unit of information-seeking actions. Step-level weighting improves BC-Sub accuracy from 11.8 to 14.5.
-
Roles of privileged information: Evidence Anchors alone improve accuracy to 14.0, while incorrect-answer feedback alone does not improve over GRPO (12.4 vs. 12.8). Combining both yields the best result (14.5). The analysis shows that Evidence Anchors primarily determine the evaluation of intermediate information-seeking actions, while incorrect-answer feedback amplifies the final-step penalty.
-
Behavioral shaping: The analysis reveals that SSPO encourages targeted, evidence-grounded information-seeking by reducing penalties for precise verification steps while amplifying penalties for broad and unfocused exploration. Steps with positive privileged-information gain use significantly fewer queries than those with negative gain.
The paper concludes that step-level self-distillation guides the agent toward more targeted and efficient information-seeking, rather than simply lifting final-answer accuracy, and demonstrates that SSPO surpasses GRPO trained for twice as many gradient steps while sustaining higher rewards and more stable gradient norms throughout optimization.
Improvements for AI systems
Improvements to AI Systems:
-
Step-Level Advantage Weighting for Long-Horizon Agents: Implement SSPO’s mechanism of converting teacher–student disagreement at each step into advantage weights within GRPO, applied only to incorrect trajectories. This decouples update direction (outcome reward) from update magnitude (step-level teacher signal), enabling stable credit assignment over dozens of steps without collapsing exploration. The improved system can train agents on sparse-reward tasks (e.g., web navigation, multi-hop QA, tool use) with significantly better sample efficiency—matching or surpassing twice the training steps of standard RL.
-
Evidence Anchors as Privileged Information for Self-Teaching: Integrate the construction of concise, step-level evidence snippets extracted from the environment (e.g., web) that capture necessary information without revealing the answer path. Use these anchors to generate privileged feedback for a self-teacher, enabling the agent to evaluate intermediate information-seeking actions (e.g., whether a search query is precise and relevant) rather than relying solely on final outcome. The improved system can learn to perform targeted, evidence-grounded exploration, reducing broad and unfocused queries while increasing precise verification steps.
-
Selective Distillation to Preserve Diversity: Adopt the policy of leaving correct trajectories untouched during training, applying step-level distillation only to incorrect ones. This prevents overfitting to the teacher’s distribution and maintains behavioral diversity (e.g., tool-use frequency), avoiding the collapse seen in direct distillation (from 17.7 to 3.5 turns). The improved system retains flexible, multi-step strategies while still benefiting from teacher guidance on mistakes.
-
Granularity-Aware Advantage Computation: Use step-level (not token-level) advantage weights, as token-level signals misalign with the natural unit of information-seeking actions. The improved system can more accurately reward or penalize entire search actions (e.g., a query, a page visit) rather than individual tokens, leading to better policy shaping (e.g., improving BC-Sub accuracy from 11.8 to 14.5).
-
Combined Privileged Signals for Dual-Purpose Feedback: Fuse Evidence Anchors (for evaluating intermediate actions) with incorrect-answer feedback (for amplifying final-step penalties) to achieve synergistic improvements. The improved system can simultaneously refine its search strategy and its final answer generation, rather than optimizing one at the expense of the other.
-
Low-Overhead Teacher Integration: Incorporate a teacher forward pass that adds only 5% to step time, making step-level self-distillation practical for real-time or large-scale training. The improved system can be deployed in resource-constrained settings without sacrificing performance gains.
Capabilities of the Improved AI System:
-
Efficient Deep Search Agents: Operates over 30+ step trajectories with sparse binary rewards, learning to navigate complex information environments (e.g., web browsing, database queries) with fewer training steps and higher final accuracy (e.g., +4.8 average score vs. +2.4 for GRPO across BrowseComp, GAIA, FRAMES).
-
Targeted Information-Seeking Behavior: Reduces unnecessary queries and focuses on evidence-grounded steps, as shown by fewer queries for steps with positive privileged-information gain. The system learns to verify precise facts rather than explore broadly, improving efficiency and interpretability.
-
Stable and Diverse Policy Optimization: Maintains high tool-use diversity and stable gradient norms throughout training, avoiding mode collapse and premature convergence, while sustaining higher reward curves than standard RL baselines.
-
Sample-Efficient Training: Achieves superior performance in 100 training steps compared to GRPO’s 200 steps, reducing computational cost and time-to-deployment for complex agentic tasks.
Abstract
Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment. On-policy self-distillation (OPSD) addresses this by using the model's own logits as dense token-level teachers, but extending it to search agents introduces a fundamental tension: the teacher, having access to privileged information such as the correct answer, produces a distribution that differs systematically from the student's exploration-based reasoning, and naive distillation causes the student to inherit this information asymmetry rather than learn better search strategies. We resolve this tension through two contributions. First, we construct Evidence Anchors, which are concise, step-level evidence snippets extracted from the web, as privileged information that captures key reasoning steps without revealing the entire answer path. Second, we propose Step-Level Self-Distilled Policy Optimization (SSPO), which converts teacher-student disagreement into step-level advantage weights within GRPO, applied exclusively to incorrect trajectories. This design decouples what to update from how much to update: the outcome reward determines the direction of policy change, while the teacher modulates its magnitude at each step. Correct trajectories are left untouched, preserving their diversity. On Qwen3-8B, SSPO consistently outperforms GRPO across BrowseComp, GAIA, and FRAMES, surpassing or matching GRPO trained with twice as many gradient steps while adding only about 5 percent overhead per step from a single additional forward pass.
Sources
- Reinforcement Learning via Self-Distillation
- An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks