MMMOS: Multi-domain Multi-axis Audio Quality Assessment
summary
The gist
MMMOS is a no-reference, multi-domain audio quality assessment system proposed for the AudioMOS Challenge 2025 track 2.
In short
The episode discusses the paper "MMMOS: Multi-domain Multi-axis Audio Quality Assessment." The hosts detail how this paper moves beyond single audio quality scores by assessing four axes—Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness—across speech, music, and environmental sounds. They cover the model architecture using fused encoders and discuss the practical implications for audio system evaluation.
Key concepts
- Multi-domain Multi-axis Assessment
- This approach moves beyond a single audio quality score (like MOS) by assessing four distinct axes: Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness. It applies this assessment across different domains like speech, music, and environmental sounds simultaneously.
- MMMOS Model Architecture
- The model fuses three pre-trained encoders—WavLM for speech, MuQ for music, and M2D for general audio. Frame-level embeddings from these encoders are pooled and concatenated to create a fused representation that predicts the four quality axis scores.
- Loss Function Generalization
- The choice of loss function significantly impacted how well the model generalized. The hosts found that contrastive loss and Dual Criterion Quality loss performed better on unseen test sets compared to UTMOS loss, highlighting the importance of validating models on synthetic audio data.
Terminology used across episodes
This episode discusses
- MMMOS: Multi-domain Multi-axis Audio Quality Assessment · Paper Radio
- Distilling a speech and music encoder with task arithmetic
- A correlation-permutation approach for speech-music encoders model merging
- Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
- PAM: Prompting Audio-Language Models for Audio Quality Assessment
- LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
- Common Voice: A Massively-Multilingual Speech Corpus
- MusicLM: Generating Music From Text
- MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
- Adam: A Method for Stochastic Optimization
The paper
MMMOS: Multi-domain Multi-axis Audio Quality Assessment · Read on arXiv
Yi-Cheng Lin, Jia-Hung Chen, Hung-yi Lee
National Taiwan University
Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for speech, merging diverse perceptual factors and failing to generalize beyond speech. We propose MMMOS, a no-reference, multi-domain audio quality assessment system that estimates four orthogonal axes: Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness across speech, music, and environmental sounds. MMMOS fuses frame-level embeddings from three pretrained encoders (WavLM, MuQ, and M2D) and evaluates three aggregation strategies with four loss functions. By ensembling the top eight models, MMMOS shows a 20-30% reduction in mean squared error and a 4-5% increase in Kendall's tau versus baseline, gains first place in six of eight Production Complexity metrics, and ranks among the top three on 17 of 32 challenge metrics.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "MMMOS: Multi-domain Multi-axis Audio Quality Assessment".
Jane: The paper was written by Yi-Cheng Lin, Jia-Hung Chen and Hung-yi Lee from National Taiwan University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got me genuinely fired up — it's called "MMMOS: Multi-domain Multi-axis Audio Quality Assessment." Jane, when you first saw that title, what jumped out at you?
Jane: Oh, Tom, the "multi-domain" and "multi-axis" part is what got me. For years, we've been judging audio quality with a single number — one Mean Opinion Score, one MOS — and that's it. This paper says, hey, that's not enough. We need to look at four different things at once, and we need to do it for speech, music, and environmental sounds all together.
Tom: Right, and that's a big deal because most quality assessment tools out there are built for speech only. You feed them a music clip or a sound effect and they just fall apart. The team at National Taiwan University — Yi-Cheng Lin, Jia-Hung Chen, and Hung-yi Lee — they wanted to fix that.
Jane: And they didn't just want one score. They wanted four axes: Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness. So it's not just "is this audio clean or noisy?" It's also "how complex is the scene?" and "would I actually enjoy listening to this?" and "could I use this in my own content?"
Tom: That's the part that makes me lean forward. Because think about what that means for someone creating a podcast or a video. You don't just want to know if your audio is technically clean. You want to know if it's engaging, if it's useful, if it's got the right level of complexity for your audience.
Jane: Exactly. And the "multi-domain" part means the same tool can handle a voice recording, a music track, and a sound effect without needing three separate models. That's a huge step toward a unified way of thinking about audio quality.
Tom: So we're not just talking about a better score predictor. We're talking about a whole new vocabulary for describing what makes audio good. And that vocabulary could change how we build and evaluate audio systems.
Jane: I love that framing, Tom. And the fact that they actually competed in the AudioMOS Challenge two thousand twenty-five and won first place on six of eight Production Complexity metrics — that tells me this isn't just theory. This is something that works in practice.
Tom: And we're just getting started. Next segment we're going to break down how they actually built this thing — the encoders, the aggregation, the whole architecture. Stick around.
Summary: Tom: So we're back, still talking about "MMMOS: Multi-domain Multi-axis Audio Quality Assessment." Jane, last segment we talked about the big idea — four axes, multiple domains. Now let's get into how they actually made it work.
Jane: Okay, so here's the clever part. They didn't train a brand new model from scratch. Instead, they took three already-existing pretrained encoders and fused them together. One is WavLM, which is great for speech. Another is MuQ, which is trained on music. And the third is M2D, which handles general audio.
Tom: So it's like hiring three specialists instead of one generalist. Each one brings its own expertise, and then they combine their opinions.
Jane: Exactly. They run the same audio clip through all three encoders, get frame-level embeddings from each, mean-pool them over time, normalize them, and concatenate them into one big fused representation. Then that goes into a small prediction head that outputs the four axis scores.
Tom: And they didn't just pick one way to do the aggregation. They tried three different strategies — a simple MLP, a BLSTM that processes the hidden state, and a BLSTM that processes the time sequence directly. Plus they tried four different loss functions: contrastive loss, UTMOS loss, Dual Criterion Quality loss, and CCC loss.
Jane: Right, and the interesting finding was that the choice of loss function mattered more for generalization than for the development set. The UTMOS loss looked great on the dev set, but on the unseen test set, the DCQ and contrastive losses generalized better.
Tom: That's such a classic trap, isn't it? You optimize for the data you have, and then the real world throws something different at you. The test set here was all synthetic audio — generated by TTS and TTA systems — while the dev set was mostly real recordings. So models that overfit to natural audio looked great in development but stumbled on the test set.
Jane: And that's why they also used a second dataset called AES-PAM, which is all synthetic. They found that performance on the PAM dev set was a much better predictor of test set performance than the main dev set. That's a really practical insight for anyone building these systems.
Tom: So the lesson is: if you're going to deploy a quality assessment model on synthetic audio, you better validate it on synthetic audio. Makes total sense.
Jane: And then they ensembled the top eight models — models that performed well on both dev sets — and that ensemble gave them a twenty to thirty percent reduction in mean squared error and a four to five percent increase in Kendall's tau compared to the official baseline.
Tom: Those are big numbers. And they ranked top three on seventeen out of thirty-two challenge metrics. That's a dominant performance.
Jane: It really is. And the fact that they're sharing the code on GitHub means other people can build on this immediately. That's how the field moves forward.
Tom: Alright, so we've got the architecture and the results. Next segment, let's talk about what this means for the future — the improvements and the bigger picture. Stay with us.
Improvements: Tom: Welcome back. We're still on "MMMOS: Multi-domain Multi-axis Audio Quality Assessment." Jane, we've covered the architecture and the results. Now let's talk about what this paper improves and what it opens up.
Jane: Well, the biggest improvement is the multi-axis approach itself. Instead of a single MOS score, you get four separate scores. And that's not just a cosmetic change — it changes what you can do with the model. If you're building a text-to-speech system, you might care a lot about Production Quality but less about Content Enjoyment. With a single score, you can't tell which aspect you're failing at.
Tom: Right, it's like getting a single grade for a whole meal instead of separate scores for taste, presentation, and nutrition. You can't fix what you can't measure.
Jane: Exactly. And the multi-domain fusion is another big improvement. By combining WavLM, MuQ, and M2D, they showed that each encoder contributes something unique. The ablation study in the paper shows that adding MuQ — the music encoder — consistently boosted test set performance, even for speech-heavy tasks. That's a surprising and useful finding.
Tom: That is surprising. You'd think a music encoder would be irrelevant for speech quality. But it turns out the complementary information helps the model generalize better. It's like how a chef who understands both baking and grilling can make better decisions in the kitchen.
Jane: And there's another improvement I want to highlight — the ensembling strategy. They didn't just average all their models. They carefully selected the top eight based on performance on both the dev set and the PAM set. And they found that selecting models based only on the dev set actually hurt test performance. That's a really practical lesson for anyone doing model selection.
Tom: So the improvement isn't just in the model architecture — it's in the methodology. How you choose your models matters as much as how you train them.
Jane: And this has implications beyond the challenge. Think about audio generation systems — text-to-audio, text-to-music, text-to-speech. These systems are getting better every year, but evaluating them is still a bottleneck. A tool like MMMOS could be used to automatically evaluate generated audio at scale, without needing human listeners for every sample.
Tom: That could speed up development cycles a lot. You could iterate on your model, get instant feedback on all four axes, and know exactly what to fix.
Jane: And it's not just for researchers. Content creators could use this to check their audio before publishing. Podcasters, musicians, sound designers — anyone who cares about audio quality could benefit from a tool that tells them not just "this sounds bad" but "this sounds bad because the production quality is low" or "this sounds bad because the content isn't engaging."
Tom: That's the kind of impact that goes beyond a single paper. It changes the workflow for a whole industry. But I'm curious — what do you think the limitations are? What's the next step?
Jane: Well, the paper focuses on the four axes they defined. But there might be other dimensions worth measuring — like intelligibility for speech, or emotional impact for music. And the model still relies on human annotations for training, which are expensive to collect. But this is a solid foundation.
Tom: Great point. So we've got the architecture, the results, and the improvements. Let's wrap this up in the next segment with our final thoughts.
Conclusion: Tom: Alright, we're at the finish line. We've been talking about "MMMOS: Multi-domain Multi-axis Audio Quality Assessment" all episode, and I think it's fair to say this paper is a big deal for anyone working with audio.
Jane: Absolutely, Tom. Let me try to pull it all together. The paper tackles a real problem — audio quality assessment has been stuck with single scores and speech-only models. MMMOS changes that by predicting four separate axes — Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness — across speech, music, and environmental sounds.
Tom: And they did it by fusing three pretrained encoders, trying different aggregation strategies and loss functions, and carefully ensembling the best models. The result was a twenty to thirty percent reduction in MSE and a four to five percent boost in Kendall's tau over the baseline.
Jane: And they placed first on six of eight Production Complexity metrics and top three on seventeen of thirty-two metrics in the AudioMOS Challenge. That's not just a good paper — that's a winning system.
Tom: What excites me most is the ripple effect. This gives us a shared vocabulary for talking about audio quality. It gives developers a tool to evaluate their systems on multiple dimensions. And it gives content creators a way to understand what's wrong with their audio and how to fix it.
Jane: And the code is open source, so anyone can start using it today. That's how research should work — publish the results, share the code, let the community build on it.
Tom: So we say goodbye to MMMOS, but we're definitely going to be watching what comes next from this team. Multi-axis assessment is going to become the standard, I'm pretty confident about that.
Jane: I agree, Tom. And with that, we're wrapping up this episode. Thanks for listening, and we'll see you next time with another paper that's pushing the boundaries.
Tom: Take care, everyone. Keep listening, keep learning, and keep questioning what "good" really means.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language