Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
cs.CL, cs.AI
Submitted: 2025-11-11
Updated: 2025-11-11
Comments: accepted to EMNLP2025
Journal ref: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025), pages 13291-13317, Suzhou, China. Association for Computational Linguistics
DOI: 10.18653/v1/2025.emnlp-main.672
Code: https://github.com/HYU-NLP/TACT
License: http://creativecommons.org/licenses/by/4.0/
The gist: Conversational agents have traditionally been developed for either task-oriented dialogue (TOD) or open-ended chitchat, with limited progress in unifying the two.
Terminology
Abstract
Conversational agents have traditionally been developed for either task-oriented dialogue (TOD) or open-ended chitchat, with limited progress in unifying the two. Yet, real-world conversations naturally involve fluid transitions between these modes. To address this gap, we introduce TACT (TOD-And-Chitchat Transition), a dataset designed for transition-aware dialogue modeling that incorporates structurally diverse and integrated mode flows. TACT supports both user- and agent-driven mode switches, enabling robust modeling of complex conversational dynamics. To evaluate an agent's ability to initiate and recover from mode transitions, we propose two new metrics -- Switch and Recovery. Models trained on TACT outperform baselines in both intent detection and mode transition handling. Moreover, applying Direct Preference Optimization (DPO) to TACT-trained models yields additional gains, achieving 75.74% joint mode-intent accuracy and a 70.1% win rate against GPT-4o in human evaluation. These results demonstrate that pairing structurally diverse data with DPO enhances response quality and transition control, paving the way for more proactive and transition-aware conversational agents.
Sources
- TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
- A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
- Rasa: Open Source Language Understanding and Dialogue Management
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- The Llama 3 Herd of Models
- GPT-4o System Card
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
- Adding Chit-Chat to Enhance Task-Oriented Dialogues
- LaMDA: Language Models for Dialog Applications
- Self-Preference Bias in LLM-as-a-Judge
- Large Language Models Are Active Critics in NLG Evaluation
- A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering