MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents
cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 23 pages, 6 figures, 5 tables. Dataset: https://huggingface.co/datasets/M2cha4l1124/MSI-Bench ; Code: https://github.com/boson-ai/MSI-Bench
Code: https://github.com/boson-ai/MSI-Bench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Speak or Stay Silent: Context-Aware Turn-Taking in Multi-Party Dialogue
- EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
- Qwen2-Audio Technical Report
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Moshi: a speech-text foundation model for real-time dialogue
- Gemma 4 Technical Report
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
- M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
- Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
- Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
- Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
- Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Qwen3-Omni Technical Report
- $\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
- IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
- Qwen3-ASR Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering