PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark

arXiv:2603.14456 · cs.CL, cs.SD · Submitted 2026-03-15 · Read on arXiv

cs.CL, cs.SD

Submitted: 2026-03-15

Updated: 2026-09-14

License: http://creativecommons.org/licenses/by/4.0/

The gist: Persian poses unique audio understanding challenges through its classical poetry, traditional music, and pervasive code-switching, none of which is captured by existing benchmarks.

Terminology

Abstract

Persian poses unique audio understanding challenges through its classical poetry, traditional music, and pervasive code-switching, none of which is captured by existing benchmarks. We introduce PARSA-Bench (Persian Audio Reasoning and Speech Assessment Benchmark), the first dedicated benchmark for evaluating LALMs on Persian language and culture. It covers 16 tasks, ten of them new, spanning speech understanding, paralinguistic analysis, and culturally grounded audio reasoning. Across most tasks, text-only baselines outperform their audio counterparts, so audio understanding rather than language knowledge remains the main limitation, and supplying the transcript alongside the audio lifts weak models to near their text-only level. The consistent exception is Persian poetry, where prosody carries information the written form cannot: audio beats text on both poetry tasks, and metre detection shows the first signs of being learnable only at the largest model scale. The dataset is publicly available at: https://huggingface.co/datasets/MohammadJRanjbar/PARSA-Bench

Sources

Related papers