MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

arXiv:2608.11755 · cs.SD, cs.CL · Submitted 2026-08-12 · Read on arXiv

Fudan University

cs.SD, cs.CL

Submitted: 2026-08-12

Updated: 2026-10-01

Code: https://github.com/WuqnEl/MuseCritic

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques This paper introduces MuseCritic, a semi-scalar reward model for the aesthetic evaluation of complete songs.

Terminology

Summary

MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

This paper introduces MuseCritic, a semi-scalar reward model for the aesthetic evaluation of complete songs. The model generates a natural-language critique covering five aesthetic dimensions and uses it as an intermediate representation to predict continuous reward scores.

The five aesthetic dimensions are: Overall Coherence, Memorability, Naturalness of Vocal Breathing and Phrasing, Clarity of Song Structure, and Overall Musicality.

MuseCritic follows a two-stage training pipeline:

  • In Stage I, a teacher model (Gemini-3-Pro) generates high-quality critiques from songs, evaluation rubrics, and expert mean ratings to construct an offline dataset (Doff), which is used for supervised fine-tuning of a pretrained audio language model (MOSS-Audio-8B-Instruct) into an SFT critique generator.

  • In Stage II, the fine-tuned model generates its own critiques for the training songs (Don), and a scalar reward head is initialized and trained using expert mean ratings as regression targets, mitigating distribution shift between training and inference.

At inference time, MuseCritic first generates a five-aspect critique and then predicts five continuous aesthetic scores conditioned on the song, rubric, and self-generated critique.

Key results:

  • On an in-domain test set of 200 SongEval songs, MuseCritic reduces macro-averaged mean squared error (MSE) from 0.2875 to 0.2316 and improves macro-averaged LCC, SRCC, and Kendall's τ to 0.9068, 0.8838, and 0.7178, respectively.

  • On the out-of-domain Music Arena benchmark with 733 preference pairs, it achieves the highest accuracy of 71.35%, which is 0.55, 2.86, and 17.60 percentage points above SongEval, Audiobox Aesthetics, and Qwen3-Omni, respectively.

  • Using MuseCritic with GRPO improves Muse-0.6B on all nine aesthetic metrics from SongEval and Audiobox Aesthetics. The largest Audiobox gains occur in production complexity (PC, +0.15) and content enjoyment (CE, +0.12), while all five SongEval dimensions improve by 0.05–0.06.

Ablation studies show:

  • Removing critiques increases macro-average MSE from 0.2316 to 0.5005, with the no-critique model concentrating predictions in a narrow range around 4.

  • Training on self-generated critiques (Don) achieves higher LCC, SRCC, and KTAU than training on offline teacher critiques (Doff) on every dimension, with macro-average MSE increasing from 0.2316 to 0.5481 when using Doff.

  • Without SFT initialization, macro-averaged MSE increases from 0.2316 to 0.7541, and macro-averaged LCC, SRCC, and KTAU decrease from 0.9068, 0.8838, and 0.7178 to 0.6695, 0.6660, and 0.5008.

The paper's contributions are:

  1. Introducing MuseCritic, a semi-scalar reward model for complete songs that generates a natural-language critique containing five dimension-specific analyses and uses this critique to predict five continuous scores.

  2. Systematically evaluating MuseCritic on in-domain SongEval ratings and out-of-domain Music Arena preferences, showing that self-generated critiques improve agreement with expert scores and mitigate critique distribution shift.

  3. Using MuseCritic as the reward model for GRPO training of Muse-0.6B, improving all nine observed metrics from SongEval and Audiobox Aesthetics.

The paper notes limitations: the critique-then-score procedure is more computationally expensive than direct score regression, and long-form song datasets pairing audio with expert aesthetic ratings remain scarce, so MuseCritic is trained primarily on Chinese and English vocal songs in SongEval.

Improvements for AI systems

Improvements to AI systems:

  1. Multi-aspect reward modeling with interpretable intermediate critiques: Replace single scalar reward models with a two-stage pipeline that first generates structured natural-language critiques across defined aesthetic dimensions (e.g., coherence, memorability, vocal naturalness, structure clarity, musicality) and then predicts continuous scores conditioned on those critiques. This reduces reward hacking and improves alignment with human aesthetic judgment by forcing the model to reason explicitly before scoring.

  2. Self-generated critique training to mitigate distribution shift: Train the reward model on critiques it generates itself (rather than only on teacher-generated critiques) to close the gap between training and inference. This improves reward accuracy and robustness when the model is used to score novel songs, as demonstrated by lower MSE and higher correlation metrics.

  3. Critique-conditioned score prediction for fine-grained feedback: Enable the reward model to output both a human-readable critique and five dimension-specific scores, allowing downstream systems to provide actionable, multi-dimensional feedback (e.g., memorability is low because the chorus lacks a distinct hook) instead of a single opaque score.

  4. RLHF/GRPO training with semi-scalar rewards: Use the multi-aspect reward model as the reward signal for policy optimization (e.g., GRPO) of a generative music model. This improves all nine aesthetic metrics simultaneously, including production complexity and content enjoyment, by optimizing for nuanced, multi-dimensional quality rather than a single aggregate score.

  5. Transferable aesthetic evaluation across domains: Apply the critique-then-score architecture to other creative domains (e.g., video, image, or text generation) where multi-dimensional aesthetic quality matters, using the same principle of intermediate natural-language reasoning to improve reward accuracy and interpretability.

What the improved AI system can do:

  • Generate music that is objectively better across multiple aesthetic dimensions (e.g., +0.15 production complexity, +0.12 content enjoyment, +0.05–0.06 on all SongEval dimensions) by optimizing against a reward model that understands and articulates why a song is good or bad.

  • Provide detailed, human-readable critiques of any song across five specific aspects, enabling artists, producers, or automated systems to identify and fix weaknesses (e.g., vocal breathing is unnatural in the second verse or song structure lacks a clear bridge).

  • Rank or compare songs more accurately than existing reward models, especially out-of-domain (e.g., 71.35% accuracy on Music Arena vs. 53.75% for Qwen3-Omni), making it useful for music recommendation, curation, and search.

  • Serve as a reliable reward signal for iterative creative generation—e.g., a music AI can generate a draft, receive a multi-aspect critique, revise based on the critique, and re-score, leading to higher-quality outputs without human intervention.

  • Adapt to new aesthetic rubrics or domains by fine-tuning the critique generator and score head on new expert ratings, while retaining the interpretable critique mechanism for transparency and debugging.

Sources

Related papers