Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents
cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- EVE-Agent: Evidence-Verifiable Self-Evolving Agents
- Self-Questioning Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- SE-Search: Self-Evolving Search Agent via Memory and Dense Reward
- CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents
- IG-Search: Step-Level Information Gain Rewards for Search-Augmented Reasoning
- Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning
- Decoupled Weight Decay Regularization
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL
- Plan Before Search: Search Agents Need Plan
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- SearchMaster: Grounded and Regulated Self-Play for Search Agents
- Knowledge-Graph Paths as Intermediate Supervision for Self-Evolving Search Agents
- Dr. Zero: Self-Evolving Search Agents without Training Data
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection