Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines
cs.CL
Submitted: 2026-09-21
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Retrieval-augmented generation (RAG) is hard to monitor in production: exhaustive relevance labels do not exist for non-stationary multi-million-passage corpora that re-index in real time.
Terminology
Abstract
Retrieval-augmented generation (RAG) is hard to monitor in production: exhaustive relevance labels do not exist for non-stationary multi-million-passage corpora that re-index in real time. As a result, retrieval quality is generally understudied and often deprioritised in favour of generation-oriented metrics. In this work, we propose auditing retrieval coverage by probing for evidence of missing documents rather than enumerating every relevant one. Our method Re:CAP (REtrieval Coverage Audit by iterative Probing) is a reference-free audit loop applied to a deployed RAG pipeline's initial answer and retrieved context: it identifies the topics already covered, generates probing questions for plausibly missing topics, retrieves candidate documents, and applies an LLM-as-judge to retain only those that introduce previously-unretrieved information. On four public benchmarks, Re:CAP recovers 9-29% of gold labels that flat BM25 top-500 cannot reach, rising to 48% on TREC-COVID. On MuSiQue Re:CAP beats flat hybrid top-500 by +12.9 pp on recall at less than half the document budget. An ensemble BM25, dense, and hybrid baseline (top-500 each) still leaves out 21.2% of gold docs on TREC-COVID that Re:CAP recovers; human annotators judge that 78.9% of those structurally distinct documents add new information to the baseline answer (Fleiss κ = 0.79, n = 123), and 73.9% on live production traffic (n = 180). End-to-end recall is reproducible to within plus or minus 1% across three independent runs, making Re:CAP a stable instrument for periodic retrieval audits.
Sources
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
- CoverageBench: Evaluating Information Coverage across Tasks and Domains
- Ragas: Automated Evaluation of Retrieval Augmented Generation
- Retrieval-Augmented Generation for Large Language Models: A Survey
- LLMs Can Patch Up Missing Relevance Judgments in Evaluation
- LANCER: LLM Reranking for Nugget Coverage
- Document Expansion by Query Prediction
- A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
- Large Language Models are not Fair Evaluators
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering