CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning

summary

Video file (mp4)

In short

The episode discusses CultureMERT, a paper that adapts an existing music model to understand Greek, Turkish, and Indian music using a two-stage continual pre-training strategy. The team also explored task arithmetic for merging single-culture models. The results show improved performance on non-Western music while maintaining performance on Western benchmarks.

Key concepts

Catastrophic Forgetting
This is a problem where a model, when retrained with new data, loses the knowledge it previously learned from old data. The paper addresses this by using a two-stage training strategy to prevent the model from forgetting its initial music knowledge when learning new musical styles.
Continual Pre-Training
This involves continually training an existing model on new data without starting from scratch. The authors used a two-stage approach: first, training lower-level feature extractors on a smaller dataset, and second, unfrozen everything to train on the full new dataset.
Task Arithmetic
This technique involves training separate models for each culture and then merging their knowledge into a single model's weight space. This method allows for combining specialized knowledge from different cultures efficiently, often performing as well as a fully trained multi-cultural model.
Stability-Plasticity Dilemma
This is the machine learning challenge of balancing a model's stability—its ability to retain old knowledge—with its plasticity—its ability to learn new things. The paper presents a practical recipe that helps achieve this balance, allowing models to adapt without forgetting prior information.

Terminology used across episodes

This episode discusses

The paper

CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning · Read on arXiv

Angelos-Nikolaos Kanatas, Charilaos Papaioannou, Alexandros Potamianos

National Technical University of Athens · Athena Research Center · Queen Mary University of London

Recent advances in music foundation models have improved audio representation learning, yet their effectiveness across diverse musical traditions remains limited. We introduce CultureMERT-95M, a multi-culturally adapted foundation model developed to enhance cross-cultural music representation learning and understanding. To achieve this, we propose a two-stage continual pre-training strategy that integrates learning rate re-warming and re-decaying, enabling stable adaptation even with limited computational resources. Training on a 650-hour multi-cultural data mix, comprising Greek, Turkish, and Indian music traditions, results in an average improvement of 4.9% in ROC-AUC and AP across diverse non-Western music auto-tagging tasks, surpassing prior state-of-the-art, with minimal forgetting on Western-centric benchmarks. We further investigate task arithmetic, an alternative approach to multi-cultural adaptation that merges single-culture adapted models in the weight space. Task arithmetic performs on par with our multi-culturally trained model on non-Western auto-tagging tasks and shows no regression on Western datasets. Cross-cultural evaluation reveals that single-culture models transfer with varying effectiveness across musical traditions, whereas the multi-culturally adapted model achieves the best overall performance. To support research on world music representation learning, we publicly release CultureMERT-95M and CultureMERT-TA-95M, fostering the development of more culturally aware music foundation models.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning".

Jane: The paper was written by Angelos-Nikolaos Kanatas, Charilaos Papaioannou and Alexandros Potamianos from National Technical University of Athens and Athena Research Center and Queen Mary University of London.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody! I'm Tom, and as always, I'm joined by the brilliant Jane. Today we are diving into a paper that just hit arXiv, and honestly, the title alone got me excited: "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning."

Jane: And I'm Jane! Tom, I have to say, this title is a mouthful, but it's tackling something really important. Most music AI models are trained almost exclusively on Western music—pop, rock, classical. But the world is full of incredible musical traditions that these models just don't understand.

Tom: Exactly! And that's the gap this paper is trying to close. The team behind this took an existing music model called MERT and adapted it to understand Greek, Turkish, and Indian music. The authors are from the National Technical University of Athens, and they've clearly put a lot of thought into this.

Jane: So, for our listeners who might not be deep in the weeds of machine learning, let me break this down. Think of MERT like a music student who has only ever listened to Western radio stations. It's really good at recognizing a guitar solo or a pop beat. But if you play it a traditional Turkish makam piece, it's completely lost.

Tom: Right, and that's a huge problem if you want to build tools for music discovery, cultural preservation, or even just recommendation systems that work for people outside the US and Europe. The paper's authors recognized this and decided to give MERT a crash course in world music.

Jane: And they didn't just throw data at it. They developed a clever two-stage training strategy to make sure the model doesn't forget what it already learned. That's a big deal because when you retrain a model on new data, it often "forgets" its old knowledge. It's like if our music student suddenly only listened to Indian classical music and then couldn't recognize a rock song anymore.

Tom: That's called catastrophic forgetting, and it's a real headache in this field. But the authors here seem to have found a way around it. I'm really curious to hear how they pulled that off, and more importantly, whether it actually works.

Jane: Me too! And I know we have our resident AI expert, Lu, joining us later to talk about the bigger picture. But for now, let's just say this paper is about making music AI more inclusive and more useful for the whole world, not just one part of it.

Tom: Well said, Jane. And speaking of the bigger picture, I think we need to get into the actual summary of what they did. That's coming up next.

Summary: Tom: So we're back, and we've got the paper "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning" on the table. Jane, you gave us the intro, but let's get into the meat of it. What did the team actually do?

Jane: Okay, so they took this pre-trained model, MERT-v1-95M, and they continually pre-trained it on a huge mix of non-Western music. We're talking about a six hundred fifty-hour dataset combining Greek traditional music, Turkish makam, and both Hindustani and Carnatic music from India.

Tom: And that's where the "continual pre-training" part comes in. They're not starting from scratch. They're taking a model that already knows a lot about music and giving it a new education. But here's the clever bit, Jane—they did it in two stages.

Jane: Right! Stage one is like a warm-up. They froze the main transformer part of the model and only trained the lower-level feature extractors on a smaller, one hundred-hour dataset. They even mixed in some Western music to keep the model grounded.

Tom: It's like stretching before a workout. You're getting the muscles ready without going full sprint. Then, in stage two, they unfroze everything and trained on the full six hundred fifty hours. This two-stage approach seems to have prevented the model from crashing and burning, which is a common problem when you dramatically shift the data distribution.

Jane: Exactly. And the results are pretty impressive. They call their main model CultureMERT, and it improved performance on non-Western music tagging tasks by an average of four point nine percent compared to the original MERT. That's a solid jump.

Tom: And they didn't just test it on the music it was trained on. They tested it on Western benchmarks too, like MagnaTagATune and FMA-medium. The model only dropped about zero point zero five percent there, which is basically negligible. So they managed to teach it new tricks without forgetting the old ones.

Jane: That's the dream scenario for continual learning. And they also explored another technique called "task arithmetic." Instead of training one big multi-cultural model, they trained separate models for each culture and then merged them together in the weight space.

Tom: Merging models in weight space sounds like science fiction, but it's actually a real technique. It's like averaging the knowledge of four specialists to get a generalist. And their merged model, CultureMERT-TA, performed just as well as the fully trained CultureMERT on non-Western tasks.

Jane: And it even did better on the Western benchmarks! That's a fantastic result because it means you don't necessarily need to do expensive multi-cultural training if you already have single-culture models lying around. You can just merge them.

Tom: So we have two paths to the same goal, and both of them beat the previous state-of-the-art on those non-Western tasks. I'm really excited to dig into the details of why the two-stage training works so well. That's up next.

Improvements: Tom: Welcome back! We're still on "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning," and Jane, we've established that the two-stage training works. But let's get into the nitty-gritty of why it's such an improvement over the naive approach.

Jane: Great question, Tom. The paper actually includes a really nice comparison table. If you just take MERT and continue training it on the new data in a single stage, it actually gets worse on the Western benchmarks. The model's ROC-AUC on MagnaTagATune drops from eighty-nine point six to eighty-six point zero. That's a big hit.

Tom: So the naive approach causes exactly the catastrophic forgetting we were worried about. But with their two-stage approach, they get the best of both worlds—improvement on non-Western music and almost no loss on Western music.

Jane: Exactly. And the reason is the stability gap. When you suddenly change the data distribution, the model's parameters get jolted, and it temporarily forgets things. Their stage one is designed to smooth out that jolt. By only updating the lower layers and using a smaller dataset with some Western replay, they let the model adjust gradually.

Tom: And they also re-warm the learning rate. That's a trick borrowed from large language model research. You don't just keep the same learning rate; you ramp it up and then decay it again. This helps the model adapt to the new domain without destroying what it already knows.

Jane: Right. And this is a huge deal for practical applications. Lu, you're our AI researcher—what do you think about this stability-plasticity balance they've achieved?

Lu: I think it's really elegant, Tom and Jane. The "stability-plasticity dilemma" is one of the oldest problems in machine learning. How do you make a model plastic enough to learn new things, but stable enough not to forget? This paper shows a practical recipe that works, and it's computationally efficient. They did all this on a single 12GB GPU, which is remarkable.

Meng: Yeah, I was going to say that. A lot of these foundation model papers require massive compute clusters. The fact that they did this on one consumer-grade GPU is a big deal for reproducibility. But I do wonder about the task arithmetic approach. How sensitive is that to the scaling factor they use?

Jane: Great question, Meng. They tested different values of the scaling factor, lambda, and found that it matters a lot. If you set it too high, like one point zero, the model's performance tanks across the board. But with a moderate value like zero point two, it works beautifully. So there's a sweet spot.

Tom: So it's not just a magic trick. You have to tune it. But once you do, you get a model that's basically as good as the fully trained one, without the extra training cost. That's a powerful tool for the community.

Meng: Absolutely. And it means you can have specialists for each culture and then combine them as needed. That's very practical.

Tom: Alright, so we've got the training strategy and the model merging. But what does this actually mean for the world? Let's bring in Lalam to give us the big-picture view in our final segment.

Conclusion: Tom: And we're back for the final stretch on "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning." Jane, we've covered the training, the results, and the merging. Let's wrap this up.

Jane: Let's do it. So to recap, the authors took a Western-centric music model and successfully adapted it to understand Greek, Turkish, and Indian music using a clever two-stage continual pre-training strategy. They also showed that you can merge single-culture models using task arithmetic to get similar results.

Tom: And both approaches beat the previous state-of-the-art on non-Western music tagging tasks. That's a clear win for the field. But Lu, what's the bigger implication here?

Lu: I think this is a stepping stone toward truly global music AI. Right now, most music foundation models are built on Western data, which means they're biased. This paper shows a practical path to fix that bias. And the fact that they're releasing the models publicly is huge for researchers in other countries who want to build on this work.

Meng: From an engineering standpoint, I love that they've shown two viable routes. If you have the data and compute, you can train a multi-cultural model from scratch. If you don't, you can train smaller single-culture models and merge them. That flexibility is really valuable.

Lalam: If I may add, the cultural impact here is profound. Music is a core part of cultural identity. By making AI models that can understand and respect diverse musical traditions, we're not just improving technology—we're helping preserve and promote cultural heritage. This model could power recommendation systems that introduce listeners to music they'd never discover otherwise, or tools that help archivists catalog and preserve endangered musical traditions.

Jane: That's a beautiful way to put it, Lalam. And it's true. This isn't just about better accuracy numbers. It's about making sure that the digital world doesn't leave entire musical cultures behind.

Tom: Well said, everyone. So, "CultureMERT" gives us a model that understands more of the world's music, and it gives us the recipe to do it again for other cultures. That's a fantastic contribution. We'll be watching to see how this develops.

Jane: Absolutely. And with that, we're wrapping up our discussion on this paper. Thanks to Lu, Meng, and Lalam for joining us, and to our listeners for tuning in. We'll be back soon with another exciting paper from the arXiv. Until then, keep listening to the world's music!

Tom: See you all next time!

More episodes

← Home