AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following
eess.AS, cs.AI, cs.SD
Submitted: 2026-06-02
Updated: 2026-09-17
Comments: EMNLP 2026
Code: https://github.com/CuCl-2/AnyAudio-Judge
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation.
Terminology
Abstract
The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on holistic scoring from general-purpose large language models, which struggle to decouple complex instructions, lack interpretability, and fail to capture fine-grained attribute mismatches. To address this, we introduce a novel dynamic rubric-based evaluation paradigm that adaptively decomposes complex audio captions into a variable number of independent, verifiable binary rubric items. To rigorously benchmark this capability, we propose the AnyAudio-Judge Bench, a comprehensive, bilingual benchmark comprising 7,920 meticulously curated samples across four diverse audio domains (speech, sound, music, and mixed), featuring deliberately constructed hard negatives. Furthermore, we construct a large-scale corpus of 105K samples with explicit Chain-of-Thought (CoT) rationales to train our dedicated evaluator, the AnyAudio-Judge model. By employing a training pipeline that combines Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO), our model successfully aligns its reasoning paths with the rubric-based scoring mechanism. Extensive experiments demonstrate that AnyAudio-Judge not only significantly enhances zero-shot alignment detection compared to state-of-the-art baselines, but also provides precise and interpretable reward signals that substantially improve instruction alignment in downstream reinforcement learning for audio generation.
Sources
- ACE-Step: A Step Towards Music Generation Foundation Model
- GPT-4 Technical Report
- MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
- Qwen3-TTS Technical Report
- PAM: Prompting Audio-Language Models for Audio Quality Assessment
- MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
- Kimi-Audio Technical Report
- InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
- DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
- AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- AudioGen: Textually Guided Audio Generation
- GSRM: Generative Speech Reward Model for Speech RLHF
- AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering
- Gemini: A Family of Highly Capable Multimodal Models
- ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
- AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
- TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions