ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems".
Jane: Recent advances in task-oriented dialogue (TOD) systems, driven by large language models (LLMs) with extensive API and tool integration, have enabled conversational agents to coordinate interleaved goals,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're starting with the title and who wrote this paper: "ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems." It’s clear they are proposing a specific tool, ATOD, to measure these advanced capabilities.
Jane: Exactly. The authors are setting up a benchmark dataset and then building an evaluation framework on top of it so that everyone can test their systems against the same rigorous standard for these complex agentic skills.
Lu: I think the core idea here is establishing a standardized way to define what "advanced" means in task-oriented dialogue, moving past simple metrics like just getting the right answer.
Meng: Standardizing evaluation is crucial because without a common yardstick, we can’t really compare how different AI architectures handle these long-term memory and coordination challenges fairly.
Lalam: It gives us a roadmap for development. Instead of guessing what makes an agent good at coordinating goals, we have a structured way to test those specific skills like dependency management and proactivity.
The paper's summary: Tom: Now let's talk about what the paper actually summarizes. Essentially, they are describing how they constructed ATOD, which is this synthetic dataset designed to mimic real-world complex interactions involving multigoal coordination and long-horizon context.
Jane: They detail a pipeline where they take existing dialogue data and use an LLM to create trajectories that explicitly show how goals interact with each other, including dependencies and the possibility of asynchronous execution.
Lu: The summary emphasizes capturing things like memory, adaptability, and proactivity because those are what traditional benchmarks often miss when looking at agents that need to maintain context across many turns.
Meng: So, the dataset isn't just random conversations; it’s engineered to force the AI to demonstrate these specific agentic traits in a controlled way. That sounds like a heavy lifting job for the generation pipeline itself.
Lalam: It’s powerful because it forces the generation process to actually simulate those tricky real-world scenarios, making sure we're testing for things like goal interleaving and dependency management right from the start.
The paper's improvements: Tom: Moving on to what they suggest as improvements, this paper proposes ATOD-Eval, which is this holistic evaluation framework that takes the ATOD dataset and translates those agentic behaviors into measurable metrics.
Jane: Instead of just looking at fluency or task success, they are proposing five key dimensions to assess: goal completion, dependency management, memory consistency, adaptability, proactivity, and multigoal coordination.
Lu: I find the structure of ATOD-Eval really interesting because it unifies evaluation by linking these abstract agentic traits directly to concrete metrics like the Dependency-Aware Goal Completion Rate and Turns to Completion.
Meng: That sounds practical; having specific numbers for how well an agent manages dependencies is much more useful for engineering teams than just a general score of success. It tells you exactly where the system is failing in its coordination logic.
Lalam: I think the idea of using an agentic memory system—combining a structured database and a semantic vector store—is key because it gives us something tangible to evaluate regarding how well the AI actually retains and updates its internal state during the conversation.
Conclusion: Tom: So, to wrap up, this paper introduces ATOD and ATOD-Eval as a unified foundation. They've provided a synthetic dataset that encodes complex agentic behaviors and a framework that allows us to systematically assess them across multiple dimensions.
Jane: The main implication is that we finally have a structured way to test if next-generation task-oriented dialogue systems can handle the complexity of real, long-running interactions involving many goals simultaneously.
Lu: It really opens up avenues for research into how memory and goal management interact in these extended scenarios, pushing us to think about context not just as a short window but as a persistent state.
Meng: For practical application, having metrics like the Dependency-Aware Goal Completion Rate gives developers clear targets for improving the coordination logic in their AI systems. It tells them precisely what needs fixing in the agent's decision-making flow.
Lalam: I think this work is significant because it provides a standardized way to push the capabilities of dialogue agents toward handling truly complex, multi-step reasoning that mimics human planning better. We’re getting closer to agents that can truly act as partners in long-horizon projects.
Amazon.com
cs.CL, cs.AI, cs.MA
Submitted: 2026-01-17
Updated: 2026-10-02
Comments: Camera-ready version accepted at AACL-IJCNLP 2026. 18 pages. Code and data: https://github.com/amazon-science/ATOD
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Recent advances in task-oriented dialogue (TOD) systems, driven by large language models (LLMs) with extensive API and tool integration, have enabled conversational agents to coordinate interleaved
Key concepts
- ATOD
- ATOD is a synthetic dataset created to test advanced task-oriented dialogue systems. It captures complex behaviors such as coordinating multiple goals, managing dependencies between those goals, and acting proactively during conversations. This dataset is built using an LLM pipeline to generate realistic, challenging multi-goal dialogues.
- ATOD-Eval
- ATOD-Eval is a comprehensive evaluation framework designed to measure the advanced agentic behaviors found in TOD systems. It assesses models across five dimensions: goal completion, dependency management, memory consistency, adaptability, proactivity, and multigoal coordination. It uses an agentic memory system to evaluate how well models maintain goal trajectories.
- Agentic Memory System
- This system is the backbone of ATOD-Eval. It consists of two stores: a structured database for symbolic metadata (Dsym) and a semantic vector store for similarity search (Dvec). The pipeline processes dialogue turns by extracting goals, checking memory, updating goal states with dependencies, and auditing for consistency.
- Dependency-Aware Goal Completion Rate (dGCR)
- dGCR is a specific metric used to measure task completion efficiency. It only counts goals that have had all their prerequisites satisfied before they are considered complete. This metric avoids bias from goals that might be blocked by unmet dependencies, providing a more accurate assessment of successful goal achievement.
Terminology
Summary
Recent advances in task-oriented dialogue (TOD) systems, driven by large language models (LLMs) with extensive API and tool integration, have enabled conversational agents to coordinate interleaved goals, maintain long-horizon context, and act proactively through asynchronous execution. This paper introduces ATOD as a benchmark dataset and ATOD-Eval as a holistic evaluation framework designed to systematically assess these advanced agentic behaviors in TOD systems.
ATOD: The Benchmark Dataset
ATOD is a synthetic dataset and generation pipeline designed to capture key characteristics of Advanced TOD, including multigoal coordination, dependency management, memory, adaptability, and proactivity.
The dataset is constructed through a modular LLM-driven pipeline involving several stages:
-
Co-occurrence Graph Construction and Goal Trajectory Sampling: This involves constructing a goal co-occurrence graph from an underlying dialogue dataset where each node represents a unique goal and weighted edges reflect empirical co-occurrence frequency. Candidate goal sets are then sampled via
stratified random walks of varying lengths over G, preserving realistic correlations while introducing diversity.
-
Annotation of Goal Trajectories and Complexity Categorization: An LLM annotates slot values,
inter-goal dependencies DS capturing prerequisite or blocking relations,
and natural-language goal descriptions. Each trajectory is assigned a complexity label based on quantitative attributes (e.g., number of goals, dependency density) and qualitative factors (e.g., interleaving or opportunities for proactivity). -
Dialogue Generation: The synthesis is conditioned on the annotated trajectory τ = (S, DS, c(S)), where the LLM produces a natural multi-turn conversation that realizes the specified goals while exhibiting
interleaving, asynchronous execution, proactive assistance, and dependency-aware coordination.
-
Turn-level Goal Status Annotation: An LLM annotator performs iterative turn-level analysis to label the status of each goal at every dialogue turn, providing a
rich reference for benchmarking multi-goal tracking and asynchronous or interleaved progressions.
ATOD-Eval: The Evaluation Framework
Building on ATOD, ATOD-Eval is proposed as a holistic evaluation framework that translates the advanced TOD dimensions into fine-grained metrics and supports reproducible offline and online evaluation. This framework unifies evaluation by jointly assessing five key dimensions: goal completion, dependency management, memory consistency, adaptability, proactivity, and multigoal coordination.
It utilizes an agentic memory system to evaluate models directly on dialogue text by assessing whether they can consistently maintain and update goal trajectories throughout interaction.
Key Evaluation Metrics
ATOD-Eval defines metrics across three dimensions:
-
Task Completion and Efficiency: This includes the
Dependency-Aware Goal Completion Rate (dGCR),
which considers only goals whose prerequisites are satisfied, avoiding bias from dependency-locked goals. It also measuresTurns to Completion (NTC)
for execution efficiency. -
Agentic Capability Metrics: These assess behaviors such as
Memory Recall Accuracy
andProactivity Effectiveness,
which evaluate goal or state changes initiated without explicit user prompts and their contextual appropriateness. -
Response Quality Metrics: These focus on turn-level relevance and dialogue-level coherence, ensuring systems maintain
natural and consistent interactions alongside effective goal management.
Agentic Memory System
The evaluation backbone incorporates an agentic memory system consisting of a dual memory store: (i) a structured goal database Dsym, which persistently records symbolic metadata like status history; and (ii) a semantic vector store Dvec, which indexes embeddings for similarity-based retrieval. The turn-level processing pipeline applies four stages to maintain the lifecycle of all goals: (i) goal extraction from the current utterance and context; (ii) existence checking against the dual memory store; (iii) updating or inserting goals with dependency evolution; and (iv) proactive auditing to keep active states consistent.
Experimental Validation
Experiments validate that ATOD-Eval enables comprehensive assessment across task completion, agentic capability, and response quality. The proposed evaluator is shown to consistently outperform competitive baselines under this evaluation setting
when assessed on goal detection and status tracking, while simultaneously achieving lower per-turn update latency and token usage.
Furthermore, correlation analysis shows that Memory Recall Accuracy correlates most strongly with dGCR in both settings,
highlighting the critical role of accurate memory in dependency-aware success. The efficiency analysis demonstrates that the method achieves the lowest per-turn update latency
and consumes fewer tokens compared to baselines like LLM-Rsum.
Conclusion
ATOD and ATOD-Eval provide a unified and scalable foundation for evaluating next-generation TOD systems.
The framework successfully addresses the gap in existing benchmarks by providing a purpose-built dataset that explicitly encodes complex agentic behaviors, enabling consistent assessment across static datasets and real-time deployments.
The gist: ATOD is a benchmark dataset and ATOD-Eval is a holistic evaluation framework designed to systematically assess advanced agentic behaviors in task-oriented dialogue systems.
Improvements for AI systems
Here are the specific improvements that can be made to current AI systems by implementing the ATOD/ATOD-Eval framework, along with what these improved systems will be able to do:
The implementation of ATOD and ATOD-Eval will fundamentally shift existing Task-Oriented Dialogue (TOD) systems from sequential, single-goal completion agents to sophisticated, agentic reasoning systems capable of managing complex, long-horizon workflows.
Here are the specific improvements and capabilities:
-
A new evaluation framework (ATOD-Eval) will be established that moves beyond simple fluency and task success rates to systematically measure advanced agentic behaviors:
-
A robust synthetic dataset (ATOD) will be created, explicitly encoding multi-goal concurrency, interleaved workflows, explicit dependencies, long-horizon memory requirements, asynchronous execution states (PENDING vs. COMPLETED), and proactive intervention opportunities.
-
An
Agentic Memory System
will be integrated into the core dialogue agent architecture using a dual memory store (symbolic metadata and semantic vector store) and a turn-level processing pipeline that dynamically extracts goals, checks against memory, updates state trajectories, and performs proactive auditing.
Specific Capabilities of the Improved AI System:
-
Maneuver Complex Goal Interleaving: The system will no longer process goals sequentially. It can simultaneously manage multiple objectives (e.g., booking a flight while concurrently arranging a hotel) and dynamically pause one goal (asynchronous execution, e.g., waiting for an external API response) to work on another, resuming seamlessly when the prerequisite is met.
-
Maintain Long-Horizon Context and State: The system will possess consistent state tracking across dialogues spanning many turns or even multiple sessions. It will accurately track the lifecycle of every goal—from OPEN (mentioned) to PENDING (active processing), COMPLETED, FAILED, or ABANDONED—ensuring that progress on one task does not corrupt the status of others.
-
Execute Dependency-Aware Workflows: The system will understand and manage prerequisites between goals (e.g.,
Payment
cannot be initiated untilBooking
is completed). It will proactively identify blocked goals and prioritize actions necessary to satisfy dependencies, ensuring logical consistency in complex, multi-step tasks. -
Demonstrate Proactive Assistance: The system will move beyond reactive responses by initiating helpful actions or reminding the user of pending tasks without being prompted (e.g.,
I see your flight is confirmed; would you like me to set a reminder to pack your passport for Sunday night?
). This demonstrates true agentic initiative. -
Achieve High Fidelity in Evaluation: By using ATOD-Eval, developers will gain fine-grained metrics (like Dependency-Aware Goal Completion Rate and Memory Recall Accuracy) that accurately reflect the performance of these advanced features, allowing for precise comparison between different LLM and memory architectures.
Abstract
Agentic task-oriented dialogue (TOD) requires systems to track concurrent goals, dependencies, and long-horizon state. We examine goal-lifecycle recovery from fixed dialogue trajectories. ATOD contains 1,000 synthetic dialogues annotated for six advanced-TOD properties, and ATOD-Eval defines metrics for dependency-sensitive completion, memory recall, and proactivity. We implement a symbolic-vector memory evaluator for ATOD-Eval. Among five prompt-only predictors and six matched-backbone memory baselines, our evaluator is the only configuration above 90% in both goal detection F1 and conditional status accuracy on medium dialogues. On complex dialogues, no baseline exceeds it on both metrics; it has the highest conditional status accuracy within the memory-based block and the lowest measured per-turn latency. These experiments assess lifecycle tracking rather than interactive agent task success. A cross-family judge swap and a manual audit with 94.0% agreement provide initial checks on measurement reliability. Code and data will be released at https://github.com/amazon-science/ATOD.
Sources
- TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
- Domain-Independent turn-level Dialogue Quality Evaluation via User Satisfaction Estimation
- MultiWOZ -- A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
- Is MultiWOZ a Solved Task? An Interactive TOD Evaluation Framework with User Simulator
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- User Simulation with Large Language Models for Evaluating Task-Oriented Dialogue
- MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs
- SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
- Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- TOAD: Task-Oriented Automatic Dialogs with Diverse Response Styles
- MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- Towards Lifelong Dialogue Agents via Timeline-based Memory Management
- RADDLE: An Evaluation Benchmark and Analysis Platform for Robust Task-oriented Dialog Systems
- LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering