EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents
Jianan Xie, Xin Sun, Zhongqi Chen, Xing Zheng, Shu Wu, Bowen Song, Liang Wang
cs.CL
Submitted: 2026-08-02
Comments: 12 pages
Code: https://github.com/JiananXie/EviSD
License: http://creativecommons.org/licenses/by/4.0/
The gist: Outcome-based reinforcement learning enables search-augmented language agents to learn from verifiable final answers, but its trajectory-level credit cannot distinguish the contributions of
Terminology
Abstract
Outcome-based reinforcement learning enables search-augmented language agents to learn from verifiable final answers, but its trajectory-level credit cannot distinguish the contributions of individual actions in a multi-turn search process. We propose EviSD, an evidence-conditioned self-distillation framework that uses instance-level supporting evidence as privileged information for search actions and golden answers as complementary privilege for answer actions. During training, the student samples actions from the original context, while the same model re-scores them as a privileged teacher under an action-aligned context. EviSD converts the detached teacher--student gap into a bounded correction to the outcome-derived GRPO advantage and applies it only to generated action spans. This design localizes privileged guidance while preserving the update direction determined by the outcome reward, without an auxiliary distillation objective or any change at inference time. Across seven question-answering benchmarks and three backbones spanning model scales and generations, EviSD achieves the highest macro-average Exact Match in all evaluated settings, outperforming the strongest compared methods by 1.3--2.3 points while modulating only 6.7%--15.1% of response tokens. Code is available at https://github.com/JiananXie/EviSD.
Sources
- What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
- Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation
- Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning
- Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
- PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning
- Self-Distilled Agentic Reinforcement Learning
- Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning
- SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning
- Group-in-Group Policy Optimization for LLM Agent Training
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
- Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
- Measuring and Narrowing the Compositionality Gap in Language Models
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Qwen2.5 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- A Survey of On-Policy Distillation for Large Language Models
- CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering