Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation
cs.SD, cs.LG
Submitted: 2026-06-24
Updated: 2026-06-24
Comments: Accepted at Interspeech 2026. Supplementary material: https://neelam472.github.io/MusicJudge/Supp.pdf (backup mirror: https://sourav-ghosh.github.io/MusicJudge/Supp.pdf )
Journal ref: Proc. Interspeech 2026, 133-137
DOI: 10.21437/Interspeech.2026-912
Project page: https://neelam472.github.io/MusicJudge/Supp.pdf
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Automatic singing quality assessment (SQA) requires evaluating lyrical correctness and musical fidelity while handling expressive variations.
Terminology
Abstract
Automatic singing quality assessment (SQA) requires evaluating lyrical correctness and musical fidelity while handling expressive variations. However, existing systems largely rely on either acoustic cues or lyric transcriptions exclusively, limiting holistic performance evaluation. Furthermore, their integration is non-trivial due to challenges in robust singing transcription amid melisma, vibrato, and tempo elasticity. To this end, we propose MusicJudge, a modality-guided framework for automated SQA that performs block-aligned multimodal analysis by coupling lyric correctness with pitch-rhythm fidelity. It detects semantically meaningful lyric blocks using multi-signal matching that integrates semantic embeddings, lexical similarity, and phonetic alignment. To improve singing audio transcription, we introduce Modality-Guided LoRA for ASR fine-tuning. Experiments across datasets demonstrate strong agreement with human expert judgments and validate the generalizability of MusicJudge.
Sources
- HCLAS-X: Hierarchical and Cascaded Lyrics Alignment System Using Multimodal Cross-Correlation
- SongTrans: An unified song transcription and alignment method for lyrics and notes
- SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment
- Automatic Estimation of Singing Voice Musical Dynamics
- SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
- gpt-oss-120b & gpt-oss-20b Model Card
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment