Prompts versus Rules: Auditing and Controlling Speech Naturalness Behaviors in Voice User Simulators
cs.CL
Submitted: 2026-10-07
Updated: 2026-10-07
Terminology
Sources
- EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
- Moshi: a speech-text foundation model for real-time dialogue
- Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
- AI Agents That Matter
- SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue
- Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
- VAmoS Bench: Voice Agent Simulation Bench
- Towards human-like spoken dialogue generation between AI agents from written dialogue
- EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions
- Evaluation and Benchmarking of LLM Agents: A Survey
- Generative Spoken Dialogue Language Modeling
- $\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
- TriageSim: A Conversational Emergency Triage Simulation Framework from Structured Electronic Health Records
- Benchmarking Complex Instruction-Following with Multiple Constraints Composition
- Following Length Constraints in Instructions
- F-Actor: Controllable Conversational Behaviour in Full-Duplex Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering