A Survey of Intent Classification and Slot-Filling Datasets for Task-Oriented Dialog
cs.CL
Submitted: 2022-07-26
Updated: 2026-09-07
Code: https://github.com/PolyAI-LDN/conversational-datasets
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Interest in dialog systems has grown substantially in the past decade.
Terminology
Abstract
Interest in dialog systems has grown substantially in the past decade. By extension, so too has interest in developing and improving intent classification and slot-filling models, which are two components that are commonly used in task-oriented dialog systems. Moreover, good evaluation benchmarks are important in helping to compare and analyze systems that incorporate such models. Unfortunately, much of the literature in the field is limited to analysis of relatively few benchmark datasets. In an effort to promote more robust analyses of task-oriented dialog systems, we have conducted a survey of publicly available datasets for the tasks of intent classification and slot-filling. We catalog the important characteristics of each dataset, and offer discussion on the applicability, strengths, and weaknesses of each. The emergence of large language models (LLMs) as capable zero-shot NLU systems gives such benchmarks a new role: the corpora cataloged here provide the principled evaluation infrastructure needed to measure LLM capabilities in structured NLU tasks, compare them against specialized models, and identify the settings---multilingual, multi-intent, low-resource---where significant gaps remain. Our goal is that this survey aids in increasing the accessibility of these datasets, which we hope will enable their use in future evaluations of intent classification and slot-filling models, whether those models are task-specific classifiers, fine-tuned language models, or zero-shot LLMs.
Sources
- SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
- Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces
- A Survey on Dialog Management: Recent Advances and Challenges
- English Machine Reading Comprehension Datasets: A Survey
- El Volumen Louder Por Favor: Code-switching in Task-oriented Semantic Parsing
- Neural Approaches to Conversational AI
- FewJoint: A Few-shot Learning Benchmark for Joint Language Understanding
- Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future
- On the Robustness of Intent Classification and Slot Labeling in Goal-oriented Dialog Systems to Real-world Noise
- Redwood: Using Collision Detection to Grow a Large-Scale Intent Classification Dataset
- DialoGLUE: A Natural Language Understanding Benchmark for Task-Oriented Dialogue
- Recent Advances in Deep Learning Based Dialogue Systems: A Systematic Survey
- Crossing the Conversational Chasm: A Primer on Natural Language Processing for Multilingual Task-Oriented Dialogue Systems
- Cross-TOP: Zero-Shot Cross-Schema Task-Oriented Parsing
- Building a Conversational Agent Overnight with Dialogue Self-Play
- Modern Question Answering Datasets and Benchmarks: A Survey
- Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing
- MultiWOZ 2.4: A Multi-Domain Task-Oriented Dialogue Dataset with Essential Annotation Corrections to Improve State Tracking Evaluation
- Meta-learning for Few-shot Natural Language Processing: A Survey
- The First Evaluation of Chinese Human-Computer Dialogue Technology
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering