Corpus2Skill: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
cs.IR, cs.AI, cs.CL, cs.MA
Submitted: 2026-04-16
Updated: 2026-08-26
Comments: Accepted to EMNLP 2026 Findings
Code: https://github.com/dukesun99/Corpus2Skill
License: http://creativecommons.org/licenses/by/4.0/
The gist: Retrieval-Augmented Generation (RAG) grounds LLM responses in external evidence but treats the model as a passive consumer of search results, with no view of how the corpus is organized or what it
Terminology
Abstract
Retrieval-Augmented Generation (RAG) grounds LLM responses in external evidence but treats the model as a passive consumer of search results, with no view of how the corpus is organized or what it has not yet seen. We present Corpus2Skill, a system-level retrieval architecture for bounded, structurally coherent corpora such as enterprise knowledge bases: an offline compiler distills the corpus into a hierarchical skill directory, and at serve time an LLM agent navigates it, drilling from a bird's-eye view through progressively finer summaries down to documents and backtracking when a branch is unproductive. On an enterprise customer-support benchmark, Corpus2Skill improves both answer quality and grounding over single-shot dense, hybrid, hierarchical-retrieval, and agentic RAG baselines at a moderate cost tradeoff, and the lead persists under encoder-matched controls and paired significance tests. An eleven-dataset study shows that corpus navigation is not a universal replacement for retrieval: it significantly wins on five datasets, ties on three, and loses on three. It helps on single-domain corpora with a recoverable topical taxonomy, but flat retrieval remains preferable on open-domain factoid pools or homogeneous-tabular corpora that defeat top-level clustering. We characterize this scope distinction as a design guideline for knowledge-grounded systems. Code is available at https://github.com/dukesun99/Corpus2Skill.
Sources
- WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation
- NaviRAG: Towards Active Knowledge Navigation for Retrieval-Augmented Generation
- A-RAG: Scaling Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- SPD-RAG: Sub-Agent Per Document Retrieval-Augmented Generation
- EvoSkill: Automated Skill Discovery for Multi-Agent Systems
- RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Retrieval-Augmented Generation for Large Language Models: A Survey
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- HAGRID: A Human-LLM Collaborative Dataset for Generative Information-Seeking with Attribution
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
- SkillX: Automatically Constructing Skill Knowledge Bases for Agents
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
- HCAG: Hierarchical Abstraction and Retrieval-Augmented Generation on Theoretical Repositories with LLMs
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG