CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval".
Jane: The paper was written by the authors from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Jane, we're looking at a heavy hitter today. The paper is CoHyDE: Iterative Co-Training of LLM Rewriter and Dense Encoder for Tool Retrieval.
Jane: It sounds like a mouthful, Tom, but the researchers at SAP Labs are tackling a massive problem for anyone building AI agents.
Tom: They're focusing on how an agent finds the right tool to use when a user's request is a bit messy or vague.
Jane: Exactly, because usually, users talk in a very casual way, while the technical API documentation is super specific and rigid.
Tom: And that mismatch is what breaks most current systems, leaving the agent lost in a sea of technical jargon.
Jane: It's like asking a friend for 'something to fix a leaky tap' and them responding with a manual for 'high-pressure hydraulic valve maintenance'.
Tom: That's a perfect analogy, Jane. The agent needs to bridge that linguistic gap.
Lu: I think this is a massive leap toward agents that can actually reason through a tool catalog rather than just searching for keywords.
Meng: I'm looking at the authors, Vaishali Senthil and the team at SAP Labs, and I wonder how this handles the scale of modern enterprise environments with thousands of APIs.
Lalam: We're seeing a move toward a more intuitive interface where the machine meets us halfway in conversation.
Tom: That's a great point, Lalam. The focus shifts toward better understanding the user's intent.
Jane: So, if we're looking at this gap between human language and technical jargon, how do they actually bridge it?
Summary: Jane: We've talked about the problem, so let's look at the actual mechanics of CoHyDE.
Tom: The core idea is this iterative co-training loop between the LLM rewriter and the dense encoder.
Jane: Think of it like a student and a teacher constantly refining their work together in a classroom.
Tom: The LLM rewriter generates these hypothetical, technical descriptions for a query, and then the encoder trains on those descriptions using InfoNCE loss.
Jane: But it doesn't stop there, does it, Tom? It's a full loop.
Tom: Exactly, Jane. The encoder's performance actually guides the rewriter's training in the next round.
Meng: They're using DPO, or Direct Preference Optimization, to align the rewriter based on the encoder's retrieval scores.
Tom: That's a clever way to make sure the rewriter is actually producing descriptions that the encoder can find easily.
Meng: It prevents the rewriter from drifting into a style that's totally useless for the search engine, which is a common failure in other methods.
Lu: It's like a dancer and a musician learning to stay in sync through constant feedback and adjustment.
Lalam: This creates a unified representation space where human intent and technical documentation finally speak the same language.
Jane: It sounds like a very tight, self-improving loop. But does it actually deliver on the promise of better retrieval?
Improvements: Tom: That's the question, Jane, and the results in the paper are quite striking.
Jane: They tested this on a subset of the ToolBench catalog, and the improvements were significant across the board.
Tom: Specifically, they saw a six point three percentage point increase in NDCG@five for those vague, conversational queries.
Jane: And don't forget the harder tiers. For the most difficult, cross-domain queries, the gains were even higher, reaching up to eight percentage points.
Tom: That's a huge jump in reliability for an agent that's trying to work in unfamiliar territory.
Meng: I was reading the ablation studies, and the rewriter warmup stage is clearly essential for this to work.
Jane: Right, Tom. You can't just start the loop with a rewriter that doesn't know the catalog's vocabulary at all.
Meng: If you skip that warmup, the encoder ends up trying to learn from descriptions that are totally disconnected from the actual tools.
Lu: It's impressive that they achieved this level of accuracy in just three rounds of training.
Lalam: The robustness they show on those hardest tiers is what really stands out to me in the data.
Tom: It really does. It proves that co-evolution is much more powerful than just trying to tune one side of the system in isolation.
Jane: It seems like the encoder and the rewriter really do need to grow together.
Conclusion: Tom: We've covered so much, Jane, with CoHyDE: Iterative Co-Training of LLM Rewriter and Dense Encoder for Tool Retrieval.
Jane: It's a brilliant piece of research that shows how we can make AI agents much more capable by letting their components learn from each other.
Tom: It's a wrap for today, but this is definitely a paper to keep an eye on as agentic workflows become the norm.
Lu: I can't wait to see how this enables agents to master entirely new domains on the fly.
Meng: From a deployment standpoint, this gives us a much more stable way to handle complex API integrations.
Lalam: This work brings us closer to a world where technology understands us as naturally as a friend does.
Jane: Thanks for listening, everyone! We'll see you at the next paper.
Tom: Bye!
cs.AI, cs.IR, cs.LG
Submitted: 2026-05-28
Updated: 2026-09-11
Comments: REALM Workshop, EMNLP 2026
Code: https://github.com/OpenBMB/ToolBench
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 83/100
The gist: This paper introduces CoHyDE, an iterative co-training framework designed to significantly improve tool retrieval performance by jointly optimizing a dense encoder and an LLM rewriter.
Key concepts
- Tool Retrieval
- This is the process of helping an AI agent find the correct technical tool or API to use when a user makes a request. It is challenging because users often speak casually while APIs require very specific jargon.
- CoHyDE
- CoHyDE stands for Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval. It is the core method that uses a continuous, self-improving loop to ensure the AI component components learn from each other.
- LLM Rewriter
- This component takes vague or conversational user queries and automatically generates hypothetical, technical descriptions. These rewritten descriptions help bridge the gap between human language and rigid API documentation.
Terminology
Summary
This paper introduces CoHyDE, an iterative co-training framework designed to significantly improve tool retrieval performance by jointly optimizing a dense encoder and an LLM rewriter. The work addresses limitations in existing systems by creating a feedback loop where the rewriter guides the encoder's understanding of tool usage, leading to robust improvements across complex retrieval tasks. This iterative approach is critical because it allows the system to bridge gaps between raw catalog information and nuanced, natural language query intent, thereby enhancing downstream agent capabilities.
The Co-Training Methodology
The core of the system involves multiple sequential training stages designed to refine both components iteratively. The process begins with foundational steps:
-
S1a involves
encoder InfoNCE warmup
over 5 epochs with a batch size of 256, requiring approximately∼3 hours
on a single H200 GPU. -
S1b focuses on the
rewriter 5-format tool-memorisation SFT
using about∼50K examples,
which takes roughly∼2 hours.
Following these initial stages, the system enters iterative co-training rounds (three total), each involving distinct modules:
-
Description Regeneration (S2): This stage regenerates descriptions over the training set (D train) at T=0, utilizing a 150-token budget via vLLM, costing
∼2 hours per round.
-
Encoder Retrain (S3): A dedicated encoder retraining pass is conducted, consuming
∼1.5 hours per round.
-
DPO Data Generation & Training (S4): This complex step involves generating data using N=4 candidates per query at T=0.7, which requires
∼6 hours per round
for generation, followed by the actual DPO training (using about 4,371 steps and LoRA r=64), costing∼4 hours per round.
Computational Scale and Resources
The computational demands are substantial, reflecting the complexity of the iterative process. The total budget for running Three rounds of co-training plus all baselines, ablations, and rejected design-choice variants
totaled approximately 400–500 GPU-hours on H200-class hardware.
For comparison, reproducing only the main result (S1a + S1b + three rounds + a single end-to-end vague evaluation) was estimated to require roughly 50 GPU-hours.
The software stack is highly specialized, relying on several key libraries and formats:
-
Encoder training utilizes an
in-house InfoNCE script built on PyTorch 2.4 and HuggingFace Transformers 4.46.
-
Rewriter SFT and DPO leverage
TRL 0.11 (SFTTrainer, DPOTrainer) with PEFT 0.13 for LoRA adapters.
-
The rewriter inference is managed using
vLLM 0.6,
and mixed-precision training employs native PyTorch bf16.
Evaluation Scope and Caveats
The evaluation framework is comprehensive, testing performance across different query tiers (G1/G2/G3) and incorporating vague paraphrases. The dataset structure includes:
-
Catalog size T = 10,000.
-
Training set D train = 104,224 (split across G1: 44,873; G2: 35,402; G3: 23,949).
-
Evaluation queries totaling
593 / 399 / 100 over G1/G2/G3 (1,092 total).
Researchers note that all reported numbers stem from a single training seed,
and while they computed paired-bootstrap 95% CIs for uncertainty, the authors caution that Multi-seed retrains, which would also bound training-side variance, are an open item.
Data Coverage and Ethical Considerations
The system operates on tools sourced from RapidAPI’s public catalog via ToolBench and ToolGen. The data coverage is limited:
-
The catalog is
entirely English language
and skewed towardconsumer-facing REST APIs (weather, sports, lifestyle, finance, entertainment).
-
Tool records contain only
API metadata (titles, endpoints, parameter schemas, free-text descriptions written by API publishers)
and do not include end-user PII. -
The authors caution that the catalog inherits selection biases—specifically
consumer-facing REST APIs over enterprise or safety-critical tools
—and that the rewriter’s behavior isencoder-specific.
Improvements for AI systems
(Internal Monologue: The provided text is a methodological appendix, not the core results section, but it contains critical limitations regarding robustness, efficiency, and generalization that must be addressed before any real-world deployment. My suggestions must focus on turning the current pipeline into a production-ready, robust system.)
Improvement: Implement Adversarial Data Augmentation and Domain Transfer Training.
-
Technical Detail: The current work relies on the ToolBench/ToolGen catalog, which exhibits selection bias (consumer-facing REST APIs). To improve generalization, the system must be trained using synthetic tool descriptions and API schemas derived from diverse, non-public domains (e.g., internal enterprise protocols, scientific instrumentation APIs) that are structurally analogous to the existing catalog but semantically orthogonal. Furthermore, incorporate adversarial examples where query paraphrases are intentionally generated to confuse the retriever (e.g., ambiguous language or domain-specific jargon outside the training distribution).
-
What it enables: The system transitions from being a specialized ToolGen retriever to a Universal API Intent Resolver. It can accurately resolve intent and retrieve relevant tools even when faced with novel, out-of-domain terminology or poor query phrasing, significantly reducing failure modes in complex enterprise environments.
The improved AI system will be a Robust, Scalable, and Responsible Universal API Intent Resolver.
It moves beyond simple retrieval by:
-
Generalizing: Handling out-of-domain and adversarial queries (Universal API Resolver).
-
Assessing Reliability: Providing confidence intervals alongside recommendations (Risk-Assessed Tool Recommendation).
-
Operating in Production: Achieving low latency via quantization and caching (High-Throughput Endpoint).
-
Ensuring Safety: Actively vetting intent and constraining dangerous calls (Responsible AI Gateway).
Abstract
Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the catalog uses technical API vocabulary that no fixed encoder can bridge on its own. The two dominant training approaches, contrastive encoder fine-tuning and HyDE-style query expansion with a frozen LLM, address this problem from opposite ends and fail in complementary directions: the fine-tuned encoder excels when the query's surface form already matches the catalog but collapses when it does not, while zero-shot HyDE is more robust to underspecified queries yet generates catalog-unaware hypothetical descriptions that degrade retrieval when queries are well-formed. We introduce CoHyDE, an iterative procedure that trains the dense encoder and the LLM rewriter as a single co-evolving system: the encoder is retrained with InfoNCE on catalog-style hypothetical descriptions produced by the rewriter, and the rewriter is preference-aligned via DPO against the encoder's retrieval scores, with both sides warm-started on the tool catalog before the loop begins. On a 10k tool subset of the ToolBench catalog, three rounds of CoHyDE improve over the strongest single-component baseline by +2.5 pp NDCG@5 on standard queries and +6.3 pp on held-out vague queries, with gains as large as +8 pp on the hardest vague tier. Ablations confirm that co-training is the key ingredient: using either component in isolation fails to match CoHyDE on both well-formed and vague queries, with losses of up to-8 pp on vague queries.
Sources
- Tool Retrieval Bridge: Aligning Vague Instructions with Retriever Preferences via Bridge Model
- ProTIP: Progressive Tool Retrieval Improves Planning
- InPars: Data Augmentation for Information Retrieval using Large Language Models
- AugTriever: Unsupervised Dense Retrieval and Domain Adaptation by Scalable Data Augmentation
- Document Expansion by Query Prediction
- Representation Learning with Contrastive Predictive Coding
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection