Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change
cs.CL, cs.AI, cs.LG, cs.SD
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/bhavik-mangla/voxparity-bench
Terminology
Sources
- RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
- NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
- Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction
- Real-Time Voice AI Hears but Does Not Listen
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- How is ChatGPT's behavior changing over time?
- ARC Prize 2024: Technical Report
- AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
- Qwen3-TTS Technical Report
- A frontend-backend architecture for tool calls in full-duplex speech models
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- Long-Tail Rebalancing for Non-Verbal Vocalization-Aware ASR: A Track 1 System for the NVVSpeech Challenge
- Learning the Difference that Makes a Difference with Counterfactually-Augmented Data
- Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
- Holistic Evaluation of Language Models
- AEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks
- Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering