CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
summary
The gist
This paper introduces CoHyDE, an iterative co-training framework designed to significantly improve tool retrieval performance by jointly optimizing a dense encoder and an LLM rewriter.
In short
The episode discusses 'CoHyDE,' a method for improving how AI agents find correct tools from large catalogs when user requests are vague or conversational. The system uses an iterative co-training loop between an LLM rewriter and a dense encoder, showing significant gains in tool retrieval accuracy.
Key concepts
- Tool Retrieval
- This is the process of helping an AI agent find the correct technical tool or API to use when a user makes a request. It is challenging because users often speak casually while APIs require very specific jargon.
- CoHyDE
- CoHyDE stands for Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval. It is the core method that uses a continuous, self-improving loop to ensure the AI component components learn from each other.
- LLM Rewriter
- This component takes vague or conversational user queries and automatically generates hypothetical, technical descriptions. These rewritten descriptions help bridge the gap between human language and rigid API documentation.
Terminology used across episodes
This episode discusses
- CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval · Paper Radio
- Tool Retrieval Bridge: Aligning Vague Instructions with Retriever Preferences via Bridge Model
- ProTIP: Progressive Tool Retrieval Improves Planning
- InPars: Data Augmentation for Information Retrieval using Large Language Models
- AugTriever: Unsupervised Dense Retrieval and Domain Adaptation by Scalable Data Augmentation
- Document Expansion by Query Prediction
- Representation Learning with Contrastive Predictive Coding
- Qwen3 Technical Report
The paper
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval · Read on arXiv
Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the catalog uses technical API vocabulary that no fixed encoder can bridge on its own. The two dominant training approaches, contrastive encoder fine-tuning and HyDE-style query expansion with a frozen LLM, address this problem from opposite ends and fail in complementary directions: the fine-tuned encoder excels when the query's surface form already matches the catalog but collapses when it does not, while zero-shot HyDE is more robust to underspecified queries yet generates catalog-unaware hypothetical descriptions that degrade retrieval when queries are well-formed. We introduce CoHyDE, an iterative procedure that trains the dense encoder and the LLM rewriter as a single co-evolving system: the encoder is retrained with InfoNCE on catalog-style hypothetical descriptions produced by the rewriter, and the rewriter is preference-aligned via DPO against the encoder's retrieval scores, with both sides warm-started on the tool catalog before the loop begins. On a 10k tool subset of the ToolBench catalog, three rounds of CoHyDE improve over the strongest single-component baseline by +2.5 pp NDCG@5 on standard queries and +6.3 pp on held-out vague queries, with gains as large as +8 pp on the hardest vague tier. Ablations confirm that co-training is the key ingredient: using either component in isolation fails to match CoHyDE on both well-formed and vague queries, with losses of up to-8 pp on vague queries.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval".
Jane: The paper was written by the authors from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Jane, we're looking at a heavy hitter today. The paper is CoHyDE: Iterative Co-Training of LLM Rewriter and Dense Encoder for Tool Retrieval.
Jane: It sounds like a mouthful, Tom, but the researchers at SAP Labs are tackling a massive problem for anyone building AI agents.
Tom: They're focusing on how an agent finds the right tool to use when a user's request is a bit messy or vague.
Jane: Exactly, because usually, users talk in a very casual way, while the technical API documentation is super specific and rigid.
Tom: And that mismatch is what breaks most current systems, leaving the agent lost in a sea of technical jargon.
Jane: It's like asking a friend for 'something to fix a leaky tap' and them responding with a manual for 'high-pressure hydraulic valve maintenance'.
Tom: That's a perfect analogy, Jane. The agent needs to bridge that linguistic gap.
Lu: I think this is a massive leap toward agents that can actually reason through a tool catalog rather than just searching for keywords.
Meng: I'm looking at the authors, Vaishali Senthil and the team at SAP Labs, and I wonder how this handles the scale of modern enterprise environments with thousands of APIs.
Lalam: We're seeing a move toward a more intuitive interface where the machine meets us halfway in conversation.
Tom: That's a great point, Lalam. The focus shifts toward better understanding the user's intent.
Jane: So, if we're looking at this gap between human language and technical jargon, how do they actually bridge it?
Summary: Jane: We've talked about the problem, so let's look at the actual mechanics of CoHyDE.
Tom: The core idea is this iterative co-training loop between the LLM rewriter and the dense encoder.
Jane: Think of it like a student and a teacher constantly refining their work together in a classroom.
Tom: The LLM rewriter generates these hypothetical, technical descriptions for a query, and then the encoder trains on those descriptions using InfoNCE loss.
Jane: But it doesn't stop there, does it, Tom? It's a full loop.
Tom: Exactly, Jane. The encoder's performance actually guides the rewriter's training in the next round.
Meng: They're using DPO, or Direct Preference Optimization, to align the rewriter based on the encoder's retrieval scores.
Tom: That's a clever way to make sure the rewriter is actually producing descriptions that the encoder can find easily.
Meng: It prevents the rewriter from drifting into a style that's totally useless for the search engine, which is a common failure in other methods.
Lu: It's like a dancer and a musician learning to stay in sync through constant feedback and adjustment.
Lalam: This creates a unified representation space where human intent and technical documentation finally speak the same language.
Jane: It sounds like a very tight, self-improving loop. But does it actually deliver on the promise of better retrieval?
Improvements: Tom: That's the question, Jane, and the results in the paper are quite striking.
Jane: They tested this on a subset of the ToolBench catalog, and the improvements were significant across the board.
Tom: Specifically, they saw a six point three percentage point increase in NDCG@five for those vague, conversational queries.
Jane: And don't forget the harder tiers. For the most difficult, cross-domain queries, the gains were even higher, reaching up to eight percentage points.
Tom: That's a huge jump in reliability for an agent that's trying to work in unfamiliar territory.
Meng: I was reading the ablation studies, and the rewriter warmup stage is clearly essential for this to work.
Jane: Right, Tom. You can't just start the loop with a rewriter that doesn't know the catalog's vocabulary at all.
Meng: If you skip that warmup, the encoder ends up trying to learn from descriptions that are totally disconnected from the actual tools.
Lu: It's impressive that they achieved this level of accuracy in just three rounds of training.
Lalam: The robustness they show on those hardest tiers is what really stands out to me in the data.
Tom: It really does. It proves that co-evolution is much more powerful than just trying to tune one side of the system in isolation.
Jane: It seems like the encoder and the rewriter really do need to grow together.
Conclusion: Tom: We've covered so much, Jane, with CoHyDE: Iterative Co-Training of LLM Rewriter and Dense Encoder for Tool Retrieval.
Jane: It's a brilliant piece of research that shows how we can make AI agents much more capable by letting their components learn from each other.
Tom: It's a wrap for today, but this is definitely a paper to keep an eye on as agentic workflows become the norm.
Lu: I can't wait to see how this enables agents to master entirely new domains on the fly.
Meng: From a deployment standpoint, this gives us a much more stable way to handle complex API integrations.
Lalam: This work brings us closer to a world where technology understands us as naturally as a friend does.
Jane: Thanks for listening, everyone! We'll see you at the next paper.
Tom: Bye!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language