EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

arXiv:2605.13841 · cs.SD, cs.AI, cs.CL, cs.LG · Submitted 2026-05-13 · Read on arXiv

cs.SD, cs.AI, cs.CL, cs.LG

Submitted: 2026-05-13

Updated: 2026-09-09

Comments: Accepted to EMNLP 2026 (Findings)

Code: https://github.com/ServiceNow/eva

Project page: https://servicenow.github.io/eva

License: http://creativecommons.org/licenses/by/4.0/

The gist: Voice agents are increasingly deployed across enterprise applications.

Terminology

Abstract

Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversation simulation and comprehensive voice-specific evaluation. We present EVA-Bench, an end-to-end evaluation framework that addresses both. On the simulation side, EVA-Bench orchestrates dynamic bot-to-bot audio conversations with automatic simulation validation that detects user simulator error and appropriately regenerates conversations before scoring. On the measurement side, EVA-Bench introduces two composite metrics: EVA-A (Accuracy) and EVA-X (Experience). EVA-Bench includes 213 scenarios across three enterprise domains, a controlled perturbation suite for accent and noise robustness, and multi-trial measurements that distinguish peak from reliable capability. Across 12 systems spanning all three architectures, we find: (1) no system simultaneously exceeds 0.5 on both EVA-A pass@1 and EVA-X pass@1; (2) peak and reliable performance diverge substantially (median pass@k--pass k gap of 0.44 on EVA-A); and (3) accent and noise perturbations expose substantial robustness gaps, with effects varying across architectures, systems, and metrics (mean Δ up to 0.314). We release EVA-Bench under an open-source license.

Sources

Related papers