SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks
summary
The gist
Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries.
In short
SEARCHSKILL teaches language models to use search tools by making query planning explicit through reusable skills. It works by having the model choose a specific search strategy before every action, rather than just searching blindly. This structured approach improves query formulation and leads to more correct answers using a small search budget.
Key concepts
- SEARCHSKILL
- A framework that makes query planning explicit by forcing the model to select a reusable search skill before executing a tool call. It introduces a token-level interface where the model declares its intended skill, like '<skill>keyword</skill>', before issuing a search or answer.
- Evolving SkillBank
- A dynamic library of search skills that grows and refines itself based on recurrent failures during training. It starts small and expands in stages, focusing first on simple skills like direct lookup before adding complex reasoning skills like 'relation-chain decomposition.'
- Select-Read-Act Protocol
- The second stage of training where the model learns to choose a skill ID from the SkillBank index and condition its next action based on that selected card's content. This teaches the model how to use the evolving skill bank as an external planning guide for tool use.
- Query Planning Diagnostics
- Metrics used to measure query quality improvement, such as 'First Query Copy' (which decreases) and 'Atomic Hop Query' (which increases). These diagnostics show that SEARCHSKILL successfully teaches models to issue more focused, less copied queries.
Terminology used across episodes
This episode discusses
- SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks · Paper Radio
- CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
- WebGPT: Browser-assisted question-answering with human feedback
- Tool Learning with Foundation Models
- Qwen2.5 Technical Report
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- Reinforcement Learning for Self-Improving Agent with Skill Library
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- ReAct: Synergizing Reasoning and Acting in Language Models
- MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
The paper
SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks · Read on arXiv
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks".
Jane: Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries.
Tom: First, who's behind it and why it matters.
Paper summary: Jane: So, to summarize our discussion on "SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks," we’ve seen that the paper proposes an explicit skill-conditioned interface for search tools and a dynamic SkillBank that learns from failure patterns to improve query planning.
Lu: The main implication is providing a reusable tool-use prior that can extend beyond just question answering, offering a structured way for models to reason through complex tasks using defined search strategies. It moves the focus from simple retrieval to intelligent query formulation.
Meng: From an engineering viewpoint, this suggests that we can train systems to be more efficient users of their search resources, leading to better practical performance on tasks that require deep knowledge synthesis across multiple documents.
Lalam: For the culture of AI development, having a mechanism where models evolve their own internal strategies based on usage patterns is valuable because it fosters a more resilient and adaptable system overall. It supports the idea that intelligence in these systems can be built iteratively through feedback.
Tom: Absolutely, so while the paper details how this two-stage supervised fine-tuning recipe works to train the model to select skills from that evolving bank, its biggest contribution is showing how we can make query planning an explicit, trainable control variable at every step. It addresses a persistent issue with broad or copied queries by giving the model a structured way to think before it hits the search engine.
Jane: That's right. The authors are demonstrating that this approach improves exact match on knowledge-intensive QA benchmarks and results in better retrieval behavior, including fewer copied first queries and more atomic hop-focused queries within a small search budget.
Lu: It’s about making the model choose a reusable search strategy before each tool call, which directly leads to better query formulation and more correct answers under limited search resources. The results show improvements in metrics like "First Query Copy (↓)," an increase in "Atomic Hop Query (↑)," and a boost in "Correct@≤ three Searches (↑)."
Meng: And they showed that this method can achieve performance close to or even stronger than large-scale reinforcement learning based search agents without the massive training costs. That means we can get sophisticated search behavior by focusing on the planning structure rather than just brute-forcing optimization through RL.
Lalam: It’s a practical way to make the underlying intelligence more reliable for users, ensuring that when an AI needs external information, it uses that information in the most effective way possible based on learned best practices. That’s where the real utility lies for widespread application.
Conclusion: Tom: So we've been digging into "SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks," and now we need to wrap up what this all means for us as listeners.
Jane: It’s a paper about how we can teach AI not just *if* they search, but *how* they decide what search query is best using these evolving skills.
Lu: Exactly, Jane; the core idea is making the query planning explicit through this skill-selection mechanism. It gives the model a structured way to approach complex information retrieval tasks.
Meng: From an engineering standpoint, it seems like a really solid way to manage the search process without just letting the AI guess at every single step in real-time.
Lalam: I think what’s truly interesting is how this framework builds its skill bank dynamically based on how the AI actually performs during training. That self-refinement is a powerful concept for long-term system development.
Tom: Right, Lalam, that dynamic adaptation is huge because it means the system learns from its own mistakes in a very structured way, which is something we need to watch closely as AI gets more complex.
Jane: And when you look at the title again, "SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks," it really lays out that the skill set isn't static; it grows and changes as the AI interacts with more difficult problems.
Lu: That evolving bank is what lets us go from simple lookups to things like complex multi-hop reasoning, which is where the real power lies for open-ended knowledge tasks.
Meng: I’m curious about the practical implications for deployment; if this skill prior works well in training, does it translate cleanly when we put these models into production systems that need reliable information access?
Lalam: For me, it means we can build AI agents that are fundamentally more reliable because they have learned a robust set of search strategies rather than just memorized specific answers. That foundational reliability really helps improve the overall culture of using AI tools responsibly.
Tom: So, we're looking at a method that formalizes the reasoning process for search, making it an explicit part of the training loop through these structured skills and a growing library of expertise.
Jane: And this moves us beyond just getting an answer; we are teaching the AI the *strategy* to get there correctly, which is a much deeper level of understanding.
Lu: The fact that they can use failure patterns to drive the skill bank evolution is fascinating; it suggests a pathway for creating more specialized and efficient search capabilities across different types of reasoning tasks.
Meng: I see it as a way to inject structured planning into the retrieval phase, which should lead to much cleaner execution paths when dealing with dense documentation.
Lalam: It’s about giving the AI a better internal map for navigating information, which could fundamentally change how we design sophisticated AI agents that need to perform research tasks autonomously.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language