LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

arXiv:2608.09934 · cs.CL, cs.AI · Submitted 2026-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "LLM Agents Factory: Retrieval of Domain-Specific LLM Agents".

Jane: The paper was written by Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh et al. from Sber AI and Moscow Institute of Physics and Technology and National University of Science and Technology MISIS and Skolkovo Institute of Science and Technology and Artificial Intelligence Research Institute.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's been making the rounds — it's called "LLM Agents Factory: Retrieval of Domain-Specific LLM Agents." And Jane, I have to say, the title alone got me excited because it promises to solve a problem I've been complaining about for months.

Jane: Oh, tell me about it, Tom. The problem here is that when you ask a large language model to act as an agent — you know, to plan, to use tools, to solve a task — the standard approach is to generate the agent's "persona" on the fly. You ask the model, "Hey, what kind of expert should handle this?" and it makes one up in real time.

Tom: And that's expensive, right? Every single user request triggers a whole agent-creation process. The authors of this paper point out that this dynamic generation is slow, it's unstable — the same request might give you a completely different agent each time — and it burns tokens like crazy.

Jane: Exactly. So what they did is they flipped the whole thing. Instead of generating an agent from scratch every time, they built a giant library of pre-made agents — over twenty thousand of them — and then they just retrieve the right one. Like picking a tool off a shelf instead of forging a new hammer every time you need to hang a picture.

Tom: That's a great analogy, Jane. And the shelf here is grounded in Wikipedia categories. They took six hundred ninety-one high-level Wikipedia domains, crossed them with forty different roles like planner, verifier, tutor, and generated profiles for each combination. So you get a "tutor for mathematics" or a "verifier for history" — all pre-made, all validated.

Jane: And the retrieval part is clever too. They use a sentence embedding model to match your query to the best agent profile. It's lightweight, it's fast, and it's deterministic — same query, same agent, every time.

Tom: Which is huge for industry, by the way. If you're building a product, you need reproducibility. You can't have your AI assistant acting like a different personality every time a user asks the same question.

Jane: Right. And the authors show that this retrieval-based approach actually beats the dynamic generation baseline on accuracy while using way fewer tokens. We're talking about a threefold reduction in token consumption and faster inference.

Tom: I love that they're not just saying "retrieval is good" — they're showing it with numbers on MMLU, BIG-bench, and BIG-bench Hard. And they even distilled the whole thing into a smaller model that can generate agents directly, which is another path entirely.

Jane: So the big idea here is: treat agent construction as a search problem, not a generation problem. And that's a shift that could really change how we build AI systems in production.

Tom: And I can't wait to dig into how they built that agent base — the filtration pipeline, the deduplication, all of it. That's coming up next.

Summary: Jane: So we're back with "LLM Agents Factory: Retrieval of Domain-Specific LLM Agents," and Tom, we teased the big idea — retrieval over generation. But let's get into the actual construction pipeline, because that's where the magic happens.

Tom: Yeah, and it's a proper engineering effort. They start with a grid of six hundred ninety-one Wikipedia domains times forty roles. For each cell, they prompt a big teacher model — GPT-OSS 120B — to generate up to five different agent profiles. Each profile has a persona, a description, and a list of tools.

Jane: And they don't just accept everything the teacher spits out. There's a three-stage filtration pipeline. First, they check syntax — the profile has to be valid JSON, and the tools have to come from a whitelist of ten executable functions.

Tom: Right, because there's no point having an agent that claims it can call a tool that doesn't exist. Then they do semantic deduplication — they embed all the profiles and use FAISS to find near-duplicates, merging them with a union-find algorithm. That removes about eighteen percent of the candidates.

Jane: And they also add a few hand-crafted "meta agents" — generic routers and triage agents — to catch queries that fall outside the domain grid. It's a safety net for out-of-domain requests.

Tom: The result is a curated base of over twenty thousand agents, each one grounded in a real Wikipedia category. That grounding is important because it means the agent's domain knowledge is tied to something structured and human-curated, not just whatever the model hallucinated.

Jane: And then the retrieval part — they serialize each agent into a flat text representation, embed it with a Sentence-BERT model, and index it with FAISS. When a query comes in, they embed it the same way and pull the top-K most similar agents.

Lu: Can I jump in here? I think the grounding in Wikipedia categories is actually the most underrated part of this paper. Most agent frameworks just let the model invent a persona out of thin air. Here, every agent is anchored to an ontological cell — a real category in human knowledge. That makes the whole system interpretable in a way that dynamic generation just isn't.

Meng: And from a practical standpoint, that grounding also means you can audit the system. If something goes wrong, you can look at which agent was retrieved and why. You can't do that if the agent was generated by a stochastic process at runtime.

Tom: Exactly. And the authors even show that retrieval-based agents beat the non-agent baseline on MMLU and BIG-bench — so it's not just cheaper, it's actually better.

Jane: But there's a trade-off, and we'll get into that when we talk about the distillation approach and the multi-agent setups. That's next.

Improvements: Tom: So we've covered the agent base and the retrieval mechanism. Now let's talk about what this paper actually improves over the state of the art. Jane, what's the biggest win here?

Jane: The biggest win is efficiency without sacrificing accuracy. They compare against AutoGen, which is a popular framework that generates agents dynamically at runtime. And on all three benchmarks — MMLU, BIG-bench, and BBH — the retrieval-based approach, which they call ALR, matches or beats AutoGen's accuracy while using about three times fewer tokens and running faster.

Meng: And that's not just a nice-to-have. In production, token cost and latency are the two things that actually determine whether you can deploy a system. If you're serving millions of requests a day, a threefold reduction in tokens is a massive cost saving.

Lu: But I want to push back on something. The paper also explores distillation — they call it ALR-Distill. They fine-tune a smaller model, Qwen3-4B, to generate agent profiles directly from queries, using synthetic training data created by the teacher model. And that approach performs on par with retrieval while doubling token consumption.

Tom: So the trade-off is interesting. Retrieval is cheaper and faster, but distillation gives you a self-contained model that doesn't need a retrieval index at runtime. For some deployments, that might be worth the extra tokens.

Jane: And then there's the multi-agent setups. They tested a few configurations where you retrieve multiple agents and let the solver pick, or where you combine retrieval with a generated agent. On the harder BBH benchmark, the multi-agent approaches actually win — they get up to sixty-nine point six percent accuracy compared to sixty-eight point five percent for single-agent retrieval.

Meng: But the multi-agent setups cost more — more tokens, more latency. So the paper is really showing a spectrum of options: cheap and fast single-agent retrieval, more expensive but more accurate multi-agent, and the distillation middle ground.

Lu: And that's the real contribution, I think. It's not just "retrieval beats generation." It's a systematic study of the trade-offs, with a concrete framework you can actually use. The authors release the code and the agent base on Hugging Face, so anyone can reproduce their results or build on top of it.

Tom: And the improvements go beyond just the benchmarks. The paper argues that retrieval-based construction improves interpretability — you can see exactly which agent was chosen and why. It improves reproducibility — same query, same agent, every time. And it reduces the token overhead from those long meta-prompts that dynamic generation requires.

Jane: So the improvements are threefold: cost, control, and consistency. And those are exactly the things that matter when you're building real systems, not just running experiments.

Meng: I'd add one more: the grounding in Wikipedia categories means the agents are aligned with human knowledge structures. That's a step toward making AI systems more trustworthy, because you can trace their behavior back to a defined domain.

Lu: And I think that's the foundation for something bigger — we'll talk about where this could go in our final segment.

Conclusion: Tom: Alright, we're wrapping up our discussion of "LLM Agents Factory: Retrieval of Domain-Specific LLM Agents." Jane, what's the one thing you want listeners to remember?

Jane: The core insight is simple: don't generate agents on the fly when you can retrieve them from a well-organized library. The authors built a repository of over twenty thousand domain-grounded agents, and they showed that retrieval beats dynamic generation on accuracy while being cheaper and faster.

Tom: And the implications are pretty big. For industry, this means AI systems can be more predictable, more auditable, and more cost-effective. For research, it opens up questions about how to build and maintain these agent libraries — how to keep them fresh, how to expand them to new domains.

Lu: I think the most exciting direction is extending this beyond Wikipedia categories. The authors mention that the same pipeline could be instantiated over medical, legal, or financial ontologies. Imagine a hospital system where every agent is grounded in a validated medical taxonomy — that would be a huge step toward trustworthy AI in healthcare.

Meng: And from an engineering perspective, I love that they released the code and the agent base. That means we can actually build on this, adapt it to our own domains, and measure the trade-offs in our own systems.

Jane: There are limitations, of course. The agent base is grounded in Wikipedia, which doesn't cover every niche domain. And the synthetic training data inherits biases from the teacher model. But the framework itself is solid, and the authors are clear about what's left to do.

Tom: Future work includes metadata-aware retrieval — searching by domain or role explicitly — and hierarchical orchestration that only invokes multi-agent reasoning when the task is complex enough to justify the cost. That's a smart way to balance quality and efficiency.

Lu: And I'd add that the distillation approach could be refined to produce more concise profiles, reducing the token overhead even further. There's a lot of room to optimize.

Tom: So we're saying goodbye to "LLM Agents Factory" — a paper that treats agent construction as a search problem and shows that sometimes the best way to build is to look it up. Thanks for listening, and we'll see you for the next paper.

Jane: Take care, everyone.

Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen Budennyy

Sber AI · Moscow Institute of Physics and Technology · National University of Science and Technology MISIS · Skolkovo Institute of Science and Technology · Artificial Intelligence Research Institute

cs.CL, cs.AI

Submitted: 2026-05-20

Updated: 2026-08-12

Comments: 7 pages, 1 figure, SIGIR 2026

DOI: 10.1145/3805712.3808515

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 64/100

The gist: "we present LLM Agents Factory, a retrieval-based framework that constructs domain-specific and Wikipedia-grounded agents on demand using a base of over 20K predetermined agent profiles." The core

Key concepts

Dynamic Generation
This is the standard approach where a large language model creates an agent's persona in real time for every user request. The authors note this method is slow, unstable (the same query might yield different agents), and consumes a lot of tokens.
Agent Base Grounding
The library of agents is not arbitrary. Each agent is tied to a specific domain, such as 'tutor for mathematics,' based on high-level Wikipedia categories. This grounding ensures the agent's knowledge is anchored in structured, human-curated information.
Retrieval Mechanism
When a query arrives, the system uses a sentence embedding model to match it against an indexed library of pre-made agents. It then pulls the top-K most similar agent profiles, making the process fast and deterministic.

Terminology

Summary

Summary

The paper introduces LLM Agents Factory, a retrieval-based framework for constructing domain-specific LLM agents on demand, addressing the limitations of dynamic (on-the-fly) agent generation in industrial settings. The authors state: we present LLM Agents Factory, a retrieval-based framework that constructs domain-specific and Wikipedia-grounded agents on demand using a base of over 20K predetermined agent profiles.

The core motivation is that while LLM agents improve task performance by decomposing problems into role-specialized behaviors, their practical deployment is limited by the computational cost and instability associated with the on-the-fly agent design for each user request. The authors identify three key drawbacks of dynamic agent generation: (1) impaired interpretability and controllability because the obtained agent profiles are not based on any task-specific ontologies or formal domain knowledge; (2) high volatility in generated entity properties because of the stochastic nature of LLM inference; and (3) dependence on long and complex meta-prompts that increase token consumption and operational costs.

The framework treats agent construction as an information retrieval (IR) problem. The paper formalizes an agent as: **a = base, domain, role, persona, description, tools **, where base is the backbone LLM, domain is the agent's expertise area, role is a discrete behavior label from a catalog of 40 roles (e.g., planner, verifier, tutor), persona and description are natural language specifications, and tools is a list of executable functions.

Agent Base Construction involves an ontology-guided process grounded in the Wikipedia category system. The authors use 691 high-level Wikipedia categories as domains and enumerate a 691×40 grid of candidate (domain, role) cells. For each cell, a teacher model (GPT-OSS 120B) generates up to 5 profiles. The construction pipeline includes: (1) Domain Choice from 691 domains; (2) Role Assignment sampled uniformly from 40 roles; (3) Profile Generation producing (persona, description) pairs; (4) Tool Selection restricted to a whitelist of 10 executable tools from the OpenAI tool-calling interface. The pipeline includes a three-stage filtration: syntax and tools checks, semantic deduplication using FAISS cosine search with union-find merging, addition of manually curated meta agents (e.g., generic router, triage, verifier), and heuristics-based filtering. This process removes approximately 18% of generated candidates.

Agent Library Retrieval (ALR) formulates agent selection as dense retrieval. Each agent profile is serialized into flat text, embedded using a Sentence-BERT bi-encoder (specifically MPNet, which consistently performs best in our preliminary experiments), and indexed with FAISS. Given a user query, the Top-K agents are retrieved based on cosine similarity: TopK(q) = arg top k h q T h a.

ALR-Distill is a supervision-based alternative that distills the agent base into a compact generator model. The training data consists of 220K (query, agent) pairs, where for each agent, 11 distinct user requests are synthesized using GPT-OSS 120B. The student model (Qwen3-4B) is fine-tuned to directly generate the JSON agent profile given a user query, maximizing P(aq). A hybrid Retrieval-Augmented Distillation variant conditions the student on retrieved Top-K profiles as context.

ALR Setups include single-agent scenarios (retrieval-based ALR and fine-tuning-based ALR-Distill) and multi-agent scenarios: ALR Top-K (retrieves Top-K agents passed to the solver) and ALR-RAG + Qwen generator (prompts an LLM with two profiles, one from ALR and one generated by Qwen).

Experiments were conducted on MMLU (N=2,070), BIG-bench (N=39,185), and BIG-bench Hard (BBH, N=2,437), using GPT-OSS 120B as the fixed solver backbone with temperature 0.3. Baselines include non-agent (no agent specification), Qwen3-4B zero-shot agent generation, and AutoGen (runtime profile generation).

Key results from Table 1:

  • ALR (single-agent) achieves 82.3% accuracy on MMLU (vs. 81.9% non-agent, 80.9% AutoGen), 85.6% on BIG-bench (vs. 84.7% non-agent, 83.4% AutoGen), and 68.5% on BBH (vs. 68.5% non-agent, 64.9% AutoGen). ALR consumes 0.8M tokens on MMLU vs. 2.4M for AutoGen, with latency 1.62s vs. 6.55s.

  • ALR-Distill achieves 82.0% on MMLU, 85.7% on BIG-bench, and 69.3% on BBH, with higher token consumption than ALR but lower than AutoGen.

  • Multi-agent setups (ALR Top-K, ALR + Qwen3-4B zero-shot, ALR + ALR-Distill) perform best on BBH (69.1%, 69.5%, 69.6% respectively) but at higher token and latency costs.

The authors conclude: "Agent Retrieval is Effective and Efficient. Compared to AutoGen, which relies on runtime profile generation, retrieval-based ALR shows consistently higher accuracy while introducing minor token and runtime overhead compared to the non-agent baseline. Specifically, our method consumes ∼3x fewer tokens across all datasets and shows faster inference."

They also note: "Retrieval vs. Fine-Tuning Trade-Off. For the single-agent setup, the fine-tuned ALR-Distill performs on par with retrieval-based ALR while doubling token consumption and marginally increasing latency. The retrieval-based approach is more favorable on our LLM Agents Factory for quality–latency balance."

Regarding single vs. multi-agent: "the relative effectiveness of single-agent and multi-agent setups depends on task complexity. On MMLU and BIG-bench, single-agent setups (ALR and ALR-Distill) achieve the highest accuracy (82.3% and 85.7%, respectively) while two-agent approaches perform better on the harder BBH dataset."

Limitations acknowledged by the authors: the evaluation isolates agent construction by fixing the solver, so results reflect efficient agent specification selection on standard question-answering and reasoning benchmarks, rather than as a complete evaluation of long-horizon, multi-turn, or tool-intensive agentic behavior. The repository is grounded in Wikipedia categories and a fixed role catalog, which can underrepresent niche industrial domains and proprietary taxonomies. Both the profile base and distillation data rely on synthetic supervision from a strong teacher model, so Biases, omissions, or hallucinated assumptions of the teacher may therefore propagate into the generated profiles and query–agent pairs.

Future work includes extending ALR with metadata-aware retrieval over explicit schema fields (domain, role, tools), hierarchical orchestration that invokes multi-agent reasoning only when task complexity justifies cost, and refining the distillation objective toward concise functional JSON profiles.

The paper concludes: "Our work addresses key limitations of runtime agent generation for industrial LLM systems: it improves interpretability through structured domain and role labels, improves reproducibility through deterministic retrieval over a validated agent base, and substantially reduces the token overhead introduced by long orchestration prompts." The implementation code and agent base are released at https://huggingface.co/frontier-ai/llm-agent-factory.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:


  1. Replace on-the-fly agent generation with retrieval from a pre-built, ontology-grounded agent repository.
  • Instead of prompting a large LLM to create a new agent profile (persona, tools, description) for every user request, I will index a static base of 20K+ validated agent profiles (each tied to a Wikipedia category and one of 40 roles) using a Sentence-BERT bi-encoder (MPNet). At inference, I retrieve the Top-K profiles via cosine similarity and inject the best one into the system prompt.
  1. Add a deterministic, multi-stage filtration pipeline for agent profiles.
  • I will enforce JSON schema validity, restrict tools to a whitelist of 10 executable functions, deduplicate semantically near-identical profiles using FAISS + union-find, and apply role-specific heuristics. This removes 18% of low-quality candidates, ensuring only executable, diverse, and interpretable agents are retrievable.
  1. Implement a hybrid retrieval-distillation mode (ALR-Distill).
  • I will fine-tune a compact model (e.g., Qwen3-4B) via LoRA on 220K synthetic (query → agent JSON) pairs generated by a teacher LLM. The student learns to directly output a valid agent profile given a user query, bypassing explicit retrieval. I will also support a retrieval-augmented variant where the student refines a retrieved Top-K profile instead of generating from scratch.
  1. Introduce a multi-agent fallback strategy for hard tasks.
  • For benchmarks like BIG-bench Hard, I will retrieve two agents (one via ALR, one via ALR-Distill) and present both to the solver, allowing it to select or combine them. This improves accuracy on complex reasoning at the cost of modestly higher token usage, but still 3x fewer tokens than AutoGen.
  1. Optimize for latency and token consumption.
  • I will cache agent embeddings and use FAISS inner-product search for sub-millisecond retrieval. I will also shorten the system prompt by using only the retrieved agent’s persona and description (not the full JSON), reducing token overhead by 3x compared to runtime generation.

  • Answer domain-specific questions with higher accuracy than non-agent baselines (e.g., +0.4% on MMLU, +0.9% on BIG-bench) while using 3x fewer tokens and 4x lower latency than dynamic agent generation (AutoGen).

  • Operate under strict industrial constraints: deterministic, interpretable agent selection (no stochastic profile generation), reproducible behavior across runs, and full auditability of which agent was used and why.

  • Handle out-of-domain or underspecified queries gracefully via manually curated meta-agents (router, triage, verifier) that are retrieved only when their score exceeds specialized profiles.

  • Scale to high-throughput production environments: the retrieval index is static, so no per-request LLM calls for agent creation; the compact distilled model can generate profiles locally without network calls.

  • Adapt to new domains by re-running the construction pipeline over a custom ontology (e.g., medical, legal, financial) without changing the retrieval or distillation logic.

  • Provide a controllable trade-off between speed and accuracy: single-agent retrieval for standard QA, multi-agent retrieval for complex reasoning, and distilled generation for fully offline or low-latency edge deployments.

Abstract

Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the on-the-fly agent design for each user request. To address this, we present LLM Agents Factory, a retrieval-based framework that constructs domain-specific and Wikipedia-grounded agents on demand using a base of over 20K predetermined agent profiles. Our framework supports two modes: (1) agent profile retrieval via semantic search and (2) distillation into a compact model fine-tuned for direct agent generation. Experiments on MMLU, BIG-bench, and BIG-bench Hard in a single-agent scenario demonstrate that our retrieval-based agent construction surpasses non-agent baselines in accuracy while matching AutoGen generation quality with a 120B backbone at a substantially lower inference cost. Our work reveals that retrieval from a structured agent repository provides a cost-efficient, accurate, and controllable alternative to dynamic agent generation, responding to the strict demands of industrial applications. We provide the implementation code and the agent base in https://huggingface.co/frontier-ai/llm-agent-factory.

Sources

Related papers