MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
cs.SD, cs.AI, cs.CL, eess.AS
Submitted: 2026-02-03
Updated: 2026-09-18
Comments: Foundation Models, Large Language Models, Native, Speech Models, Arabic, AI-persona, Persona-conditioned-conversations
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Audio large language models (AudioLLMs) enable instruction following over speech and general audio, but progress is limited by the scarcity of diverse, conversational, and instruction-aligned
Terminology
Abstract
Audio large language models (AudioLLMs) enable instruction following over speech and general audio, but progress is limited by the scarcity of diverse, conversational, and instruction-aligned speech--text data. This gap is particularly pronounced for persona-grounded and dialectal interactions, where collecting real multi-speaker recordings remains costly and slow. We introduce MENASpeechBank, a reference speech bank comprising 18K high-quality utterances from 124 speakers spanning multiple MENA countries, covering English, Modern Standard Arabic (MSA), and regional Arabic varieties. We develop a controllable data pipeline that (i) constructs persona profiles enriched with World Values Survey (WVS) inspired attributes, (ii) defines a taxonomy driven 5Kconversational scenarios, (iii) matches personas to scenarios via semantic similarity, (iv) generates 417K role-play conversations with an LLM where the user speaks as the persona and the assistant behaves as a helpful agent, and (v) produces speaker-conditioned user-turn audio (synthetic) from reference recordings to preserve speaker diversity. We evaluate synthetic and human recorded conversations and provide an analysis. We will make the MENASpeechBank available for the community.(https://huggingface.co/datasets/QCRI/MenaSpeechBank)
Sources
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- NativQA: Multilingual Culturally-Aligned Natural Query for LLMs
- Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- The optimization of paths in the $R^{3,1}$ space time by Markov Chain Monte Carlos
- SpeechBrain: A General-Purpose Speech Toolkit
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment