Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems
cs.IR, cs.CL
Submitted: 2026-09-08
Updated: 2026-09-28
Code: https://github.com/quickwitoss/tantivy
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries
Terminology
Abstract
Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant documents per query. No existing public benchmark evaluates this setting: large-scale collections typically provide only a small number of evaluation queries, whereas benchmarks with many queries generally contain only millions of documents. Moreover, most benchmarks assess human-written queries, while the first-stage retrievers in agentic RAG pipelines serve machine-written reformulations whose distribution differs from human search behavior. To overcome these evaluation gaps, we introduce Q2D-Web (Query2Doc-Web), a large-scale agentic retrieval benchmark consisting of a 190M-document web corpus and 70k agentic search queries in ten languages, reformulated from real-world user queries in production systems. Q2D-Web provides three sets of fixed relevance judgments: agent citations, production rankings, and a combined set that unions both signals and adds LLM-based judgments of unlabeled pooled documents to reduce false negatives. We benchmark 13 retrievers including lexical, dense, and late-interaction models and find that their relative ordering is largely insensitive to the choice of judgment set, while diverging substantially across topical domains, query languages, and query types. To enable fast evaluation, we also study subcorpus sampling as an approximation to full-corpus evaluations. Retaining a third of the corpus, selected by reciprocal rank fusion over pooled retriever runs, preserves the full-corpus model ranking under the combined judgments while raising absolute Recall@1000 only by 3 to 7 points. The public leaderboard is accessible under: https://huggingface.co/spaces/perplexity-ai/q2d-web-leaderboard
Sources
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
- A Survey on LLM-as-a-Judge
- AnglE-optimized Text Embeddings
- Multi-Stage Document Ranking with BERT
- ClueWeb22: 10 Billion Web Documents with Visual and Semantic Information
- Ragnar\"ok: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
- DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
- EmbeddingGemma: Powerful and Lightweight Text Representations
- Multilingual E5 Text Embeddings: A Technical Report
- Measuring short-form factuality in large language models
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- Arctic-Embed 2.0: Multilingual Retrieval Without Compromise
- Jasper and Stella: distillation of SOTA embedding models
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG