EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
cs.SD, cs.AI, cs.CL, cs.LG
Submitted: 2026-05-13
Updated: 2026-09-09
Comments: Accepted to EMNLP 2026 (Findings)
Code: https://github.com/ServiceNow/eva
Project page: https://servicenow.github.io/eva
License: http://creativecommons.org/licenses/by/4.0/
The gist: Voice agents are increasingly deployed across enterprise applications.
Terminology
Abstract
Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversation simulation and comprehensive voice-specific evaluation. We present EVA-Bench, an end-to-end evaluation framework that addresses both. On the simulation side, EVA-Bench orchestrates dynamic bot-to-bot audio conversations with automatic simulation validation that detects user simulator error and appropriately regenerates conversations before scoring. On the measurement side, EVA-Bench introduces two composite metrics: EVA-A (Accuracy) and EVA-X (Experience). EVA-Bench includes 213 scenarios across three enterprise domains, a controlled perturbation suite for accent and noise robustness, and multi-trial measurements that distinguish peak from reliable capability. Across 12 systems spanning all three architectures, we find: (1) no system simultaneously exceeds 0.5 on both EVA-A pass@1 and EVA-X pass@1; (2) peak and reliable performance diverge substantially (median pass@k--pass k gap of 0.44 on EVA-A); and (3) accent and noise perturbations expose substantial robustness gaps, with effects varying across architectures, systems, and metrics (mean Δ up to 0.314). We release EVA-Bench under an open-source license.
Sources
- Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms
- Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
- Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
- Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
- SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- $\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
- Designing Sequence with Minimum PSL Using Chebyshev Distance and its Application for Chaotic MIMO Radar Waveform Design
- Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment