Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
cs.CL, cs.AI, cs.SD
Submitted: 2026-08-18
Updated: 2026-09-18
Comments: Multi-turn Conversational AI; Multimodal Dialogue; AudioLLMs; Conversational Memory; Tool-Augmented Agents; Dialogue Evaluation
Code: https://github.com/faiza-sfa/multiturn-conversational-ai-survey
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction.
Terminology
Abstract
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
Sources
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
- Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
- mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
- FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Qwen2-Audio Technical Report
- MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Moshi: a speech-text foundation model for real-time dialogue
- MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
- MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- PIPPA: A Partially Synthetic Conversational Dataset
- TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
- M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
- GPT-4o System Card
- ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering