ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding

arXiv:2609.22771 · cs.SD, cs.AI, cs.LG, eess.AS, eess.SP · Submitted 2026-09-19 · Read on arXiv

cs.SD, cs.AI, cs.LG, eess.AS, eess.SP

Submitted: 2026-09-19

Updated: 2026-09-19

Comments: Accepted to Interspeech 2026. Project Website: https://nishitanand.github.io/paralinguistic-understanding-llm/

Project page: https://nishitanand.github.io/paralinguistic-understanding-llm

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

The gist: Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and

Terminology

Abstract

Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristics and create a dataset of over 1.2M Audio-QA pairs. We develop ParA-LLM, trained with a two-stage curriculum: first on single-attribute questions to build foundational knowledge, then on multi-attribute questions for joint reasoning over speaker and acoustic characteristics. We also release ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories, where frontier models like GPT-4o-Audio achieve only 36% accuracy. ParA-LLM surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.

Related papers