MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

arXiv:2602.07036 · cs.SD, cs.AI, cs.CL, eess.AS · Submitted 2026-02-03 · Read on arXiv

cs.SD, cs.AI, cs.CL, eess.AS

Submitted: 2026-02-03

Updated: 2026-09-18

Comments: Foundation Models, Large Language Models, Native, Speech Models, Arabic, AI-persona, Persona-conditioned-conversations

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

The gist: Audio large language models (AudioLLMs) enable instruction following over speech and general audio, but progress is limited by the scarcity of diverse, conversational, and instruction-aligned

Terminology

Abstract

Audio large language models (AudioLLMs) enable instruction following over speech and general audio, but progress is limited by the scarcity of diverse, conversational, and instruction-aligned speech--text data. This gap is particularly pronounced for persona-grounded and dialectal interactions, where collecting real multi-speaker recordings remains costly and slow. We introduce MENASpeechBank, a reference speech bank comprising 18K high-quality utterances from 124 speakers spanning multiple MENA countries, covering English, Modern Standard Arabic (MSA), and regional Arabic varieties. We develop a controllable data pipeline that (i) constructs persona profiles enriched with World Values Survey (WVS) inspired attributes, (ii) defines a taxonomy driven 5Kconversational scenarios, (iii) matches personas to scenarios via semantic similarity, (iv) generates 417K role-play conversations with an LLM where the user speaks as the persona and the assistant behaves as a helpful agent, and (v) produces speaker-conditioned user-turn audio (synthetic) from reference recordings to preserve speaker diversity. We evaluate synthetic and human recorded conversations and provide an analysis. We will make the MENASpeechBank available for the community.(https://huggingface.co/datasets/QCRI/MenaSpeechBank)

Sources

Related papers