RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation
cs.IR, cs.AI, cs.CL
Submitted: 2026-01-30
Updated: 2026-01-30
Comments: Accepted for publication at CAIN 2026 (5th International Conference on AI Engineering)
Code: https://github.com/lorenzbrehme/RAG-DIVE
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
- Unsupervised Evaluation of Interactive Dialog with DialoGPT
- LLM Evaluators Recognize and Favor Their Own Generations
- Know What You Don't Know: Unanswerable Questions for SQuAD
- CLAPNQ: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems
- ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
- Evaluating Retrieval Quality in Retrieval-Augmented Generation
- ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate Disclosures
- RACE: Retrieval-Augmented Commit Message Generation
- MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
- MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries
- Searching for Best Practices in Retrieval-Augmented Generation
- CRAG -- Comprehensive RAG Benchmark
- A Comprehensive Assessment of Dialog Evaluation Metrics
- Multi-Source Knowledge Pruning for Retrieval-Augmented Generation: A Benchmark and Empirical Study
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
- Non-Determinism of "Deterministic" LLM Settings
- Retrieval-Augmented Generation in Industry: An Interview Study on Use Cases, Requirements, Challenges, and Evaluation
- Out of Style: RAG's Fragility to Linguistic Variation
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG