CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning

arXiv:2506.17818 · cs.SD, cs.AI, cs.LG, eess.AS · Submitted 2025-06-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning".

Jane: The paper was written by Angelos-Nikolaos Kanatas, Charilaos Papaioannou and Alexandros Potamianos from National Technical University of Athens and Athena Research Center and Queen Mary University of London.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody! I'm Tom, and as always, I'm joined by the brilliant Jane. Today we are diving into a paper that just hit arXiv, and honestly, the title alone got me excited: "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning."

Jane: And I'm Jane! Tom, I have to say, this title is a mouthful, but it's tackling something really important. Most music AI models are trained almost exclusively on Western music—pop, rock, classical. But the world is full of incredible musical traditions that these models just don't understand.

Tom: Exactly! And that's the gap this paper is trying to close. The team behind this took an existing music model called MERT and adapted it to understand Greek, Turkish, and Indian music. The authors are from the National Technical University of Athens, and they've clearly put a lot of thought into this.

Jane: So, for our listeners who might not be deep in the weeds of machine learning, let me break this down. Think of MERT like a music student who has only ever listened to Western radio stations. It's really good at recognizing a guitar solo or a pop beat. But if you play it a traditional Turkish makam piece, it's completely lost.

Tom: Right, and that's a huge problem if you want to build tools for music discovery, cultural preservation, or even just recommendation systems that work for people outside the US and Europe. The paper's authors recognized this and decided to give MERT a crash course in world music.

Jane: And they didn't just throw data at it. They developed a clever two-stage training strategy to make sure the model doesn't forget what it already learned. That's a big deal because when you retrain a model on new data, it often "forgets" its old knowledge. It's like if our music student suddenly only listened to Indian classical music and then couldn't recognize a rock song anymore.

Tom: That's called catastrophic forgetting, and it's a real headache in this field. But the authors here seem to have found a way around it. I'm really curious to hear how they pulled that off, and more importantly, whether it actually works.

Jane: Me too! And I know we have our resident AI expert, Lu, joining us later to talk about the bigger picture. But for now, let's just say this paper is about making music AI more inclusive and more useful for the whole world, not just one part of it.

Tom: Well said, Jane. And speaking of the bigger picture, I think we need to get into the actual summary of what they did. That's coming up next.

Summary: Tom: So we're back, and we've got the paper "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning" on the table. Jane, you gave us the intro, but let's get into the meat of it. What did the team actually do?

Jane: Okay, so they took this pre-trained model, MERT-v1-95M, and they continually pre-trained it on a huge mix of non-Western music. We're talking about a six hundred fifty-hour dataset combining Greek traditional music, Turkish makam, and both Hindustani and Carnatic music from India.

Tom: And that's where the "continual pre-training" part comes in. They're not starting from scratch. They're taking a model that already knows a lot about music and giving it a new education. But here's the clever bit, Jane—they did it in two stages.

Jane: Right! Stage one is like a warm-up. They froze the main transformer part of the model and only trained the lower-level feature extractors on a smaller, one hundred-hour dataset. They even mixed in some Western music to keep the model grounded.

Tom: It's like stretching before a workout. You're getting the muscles ready without going full sprint. Then, in stage two, they unfroze everything and trained on the full six hundred fifty hours. This two-stage approach seems to have prevented the model from crashing and burning, which is a common problem when you dramatically shift the data distribution.

Jane: Exactly. And the results are pretty impressive. They call their main model CultureMERT, and it improved performance on non-Western music tagging tasks by an average of four point nine percent compared to the original MERT. That's a solid jump.

Tom: And they didn't just test it on the music it was trained on. They tested it on Western benchmarks too, like MagnaTagATune and FMA-medium. The model only dropped about zero point zero five percent there, which is basically negligible. So they managed to teach it new tricks without forgetting the old ones.

Jane: That's the dream scenario for continual learning. And they also explored another technique called "task arithmetic." Instead of training one big multi-cultural model, they trained separate models for each culture and then merged them together in the weight space.

Tom: Merging models in weight space sounds like science fiction, but it's actually a real technique. It's like averaging the knowledge of four specialists to get a generalist. And their merged model, CultureMERT-TA, performed just as well as the fully trained CultureMERT on non-Western tasks.

Jane: And it even did better on the Western benchmarks! That's a fantastic result because it means you don't necessarily need to do expensive multi-cultural training if you already have single-culture models lying around. You can just merge them.

Tom: So we have two paths to the same goal, and both of them beat the previous state-of-the-art on those non-Western tasks. I'm really excited to dig into the details of why the two-stage training works so well. That's up next.

Improvements: Tom: Welcome back! We're still on "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning," and Jane, we've established that the two-stage training works. But let's get into the nitty-gritty of why it's such an improvement over the naive approach.

Jane: Great question, Tom. The paper actually includes a really nice comparison table. If you just take MERT and continue training it on the new data in a single stage, it actually gets worse on the Western benchmarks. The model's ROC-AUC on MagnaTagATune drops from eighty-nine point six to eighty-six point zero. That's a big hit.

Tom: So the naive approach causes exactly the catastrophic forgetting we were worried about. But with their two-stage approach, they get the best of both worlds—improvement on non-Western music and almost no loss on Western music.

Jane: Exactly. And the reason is the stability gap. When you suddenly change the data distribution, the model's parameters get jolted, and it temporarily forgets things. Their stage one is designed to smooth out that jolt. By only updating the lower layers and using a smaller dataset with some Western replay, they let the model adjust gradually.

Tom: And they also re-warm the learning rate. That's a trick borrowed from large language model research. You don't just keep the same learning rate; you ramp it up and then decay it again. This helps the model adapt to the new domain without destroying what it already knows.

Jane: Right. And this is a huge deal for practical applications. Lu, you're our AI researcher—what do you think about this stability-plasticity balance they've achieved?

Lu: I think it's really elegant, Tom and Jane. The "stability-plasticity dilemma" is one of the oldest problems in machine learning. How do you make a model plastic enough to learn new things, but stable enough not to forget? This paper shows a practical recipe that works, and it's computationally efficient. They did all this on a single 12GB GPU, which is remarkable.

Meng: Yeah, I was going to say that. A lot of these foundation model papers require massive compute clusters. The fact that they did this on one consumer-grade GPU is a big deal for reproducibility. But I do wonder about the task arithmetic approach. How sensitive is that to the scaling factor they use?

Jane: Great question, Meng. They tested different values of the scaling factor, lambda, and found that it matters a lot. If you set it too high, like one point zero, the model's performance tanks across the board. But with a moderate value like zero point two, it works beautifully. So there's a sweet spot.

Tom: So it's not just a magic trick. You have to tune it. But once you do, you get a model that's basically as good as the fully trained one, without the extra training cost. That's a powerful tool for the community.

Meng: Absolutely. And it means you can have specialists for each culture and then combine them as needed. That's very practical.

Tom: Alright, so we've got the training strategy and the model merging. But what does this actually mean for the world? Let's bring in Lalam to give us the big-picture view in our final segment.

Conclusion: Tom: And we're back for the final stretch on "CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning." Jane, we've covered the training, the results, and the merging. Let's wrap this up.

Jane: Let's do it. So to recap, the authors took a Western-centric music model and successfully adapted it to understand Greek, Turkish, and Indian music using a clever two-stage continual pre-training strategy. They also showed that you can merge single-culture models using task arithmetic to get similar results.

Tom: And both approaches beat the previous state-of-the-art on non-Western music tagging tasks. That's a clear win for the field. But Lu, what's the bigger implication here?

Lu: I think this is a stepping stone toward truly global music AI. Right now, most music foundation models are built on Western data, which means they're biased. This paper shows a practical path to fix that bias. And the fact that they're releasing the models publicly is huge for researchers in other countries who want to build on this work.

Meng: From an engineering standpoint, I love that they've shown two viable routes. If you have the data and compute, you can train a multi-cultural model from scratch. If you don't, you can train smaller single-culture models and merge them. That flexibility is really valuable.

Lalam: If I may add, the cultural impact here is profound. Music is a core part of cultural identity. By making AI models that can understand and respect diverse musical traditions, we're not just improving technology—we're helping preserve and promote cultural heritage. This model could power recommendation systems that introduce listeners to music they'd never discover otherwise, or tools that help archivists catalog and preserve endangered musical traditions.

Jane: That's a beautiful way to put it, Lalam. And it's true. This isn't just about better accuracy numbers. It's about making sure that the digital world doesn't leave entire musical cultures behind.

Tom: Well said, everyone. So, "CultureMERT" gives us a model that understands more of the world's music, and it gives us the recipe to do it again for other cultures. That's a fantastic contribution. We'll be watching to see how this develops.

Jane: Absolutely. And with that, we're wrapping up our discussion on this paper. Thanks to Lu, Meng, and Lalam for joining us, and to our listeners for tuning in. We'll be back soon with another exciting paper from the arXiv. Until then, keep listening to the world's music!

Tom: See you all next time!

Angelos-Nikolaos Kanatas, Charilaos Papaioannou, Alexandros Potamianos

National Technical University of Athens · Athena Research Center · Queen Mary University of London

cs.SD, cs.AI, cs.LG, eess.AS

Submitted: 2025-06-21

Updated: 2026-08-18

Comments: 10 pages, 4 figures, accepted to the 26th International Society for Music Information Retrieval conference (ISMIR 2025), to be held in Daejeon, South Korea

Code: https://github.com/facebookresearch/fairseq

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

Key concepts

Catastrophic Forgetting
This is a problem where a model, when retrained with new data, loses the knowledge it previously learned from old data. The paper addresses this by using a two-stage training strategy to prevent the model from forgetting its initial music knowledge when learning new musical styles.
Continual Pre-Training
This involves continually training an existing model on new data without starting from scratch. The authors used a two-stage approach: first, training lower-level feature extractors on a smaller dataset, and second, unfrozen everything to train on the full new dataset.
Task Arithmetic
This technique involves training separate models for each culture and then merging their knowledge into a single model's weight space. This method allows for combining specialized knowledge from different cultures efficiently, often performing as well as a fully trained multi-cultural model.
Stability-Plasticity Dilemma
This is the machine learning challenge of balancing a model's stability—its ability to retain old knowledge—with its plasticity—its ability to learn new things. The paper presents a practical recipe that helps achieve this balance, allowing models to adapt without forgetting prior information.

Terminology

Summary

Summary

This paper introduces CultureMERT-95M, a multi-culturally adapted music foundation model developed to enhance cross-cultural music representation learning and understanding. The authors propose a two-stage continual pre-training strategy that integrates learning rate re-warming and re-decaying, enabling stable adaptation even with limited computational resources. Training on a 650-hour multi-cultural data mix, comprising Greek, Turkish, and Indian music traditions, results in an average improvement of 4.9% in ROC-AUC and AP across diverse non-Western music auto-tagging tasks, surpassing prior state-of-the-art, with minimal forgetting on Western-centric benchmarks.

The authors further investigate task arithmetic, an alternative approach to multi-cultural adaptation that merges single-culture adapted models in the weight space. Task arithmetic performs on par with the multi-culturally trained model on non-Western auto-tagging tasks and shows no regression on Western datasets. Cross-cultural evaluation reveals that single-culture models transfer with varying effectiveness across musical traditions, whereas the multi-culturally adapted model achieves the best overall performance. To support research on world music representation learning, the authors publicly release CultureMERT-95M and CultureMERT-TA-95M.

The paper notes that most existing foundation models for music have been trained primarily on Western-centric datasets, limiting their ability to represent diverse musical styles. Many musical traditions, including Turkish, Indian, and Greek traditional music, feature unique melodic structures, modal or tonal systems, and rhythmic patterns that are not adequately captured by these models. Continual pre-training (CPT) has emerged as an effective approach in large language models and multimodal learning, enabling models to incrementally adapt to new domains without full re-training. Model merging, particularly task arithmetic, has also proven effective for adapting pre-trained models across multiple domains by combining domain-specific parameters in weight space.

The main contributions are: (1) the first study to explore continual pre-training and task arithmetic for cross-cultural adaptation in MIR; (2) a two-stage CPT strategy that stabilizes training, mitigates catastrophic forgetting, and facilitates effective adaptation under constrained computational resources; (3) CultureMERT outperforms the original MERT-v1 by an average of 4.9% across ROC-AUC and AP on culturally diverse non-Western music tagging tasks with minimal forgetting on Western benchmarks; (4) culturally adapted models surpass previous state-of-the-art results across all evaluated non-Western music tagging tasks; (5) analysis of cross-cultural transferability showing single-culture adaptations exhibit varying degrees of transfer across cultural domains.

For datasets, the authors use MagnaTagATune (MTAT) and FMA-medium to represent Western music, and for non-Western traditions, they incorporate the Lyra corpus featuring Greek traditional and folk music, along with three collections from the CompMusic Corpora: Turkish-makam, Hindustani, and Carnatic music. They extract 200 hours each from Turkish-makam, Carnatic, and Hindustani datasets, and 50 hours from Lyra, combining these into a unified 650-hour dataset for multi-cultural continual pre-training.

The continual pre-training objective follows the self-supervised masked language modeling objective of MERTRVQ-VAE, where two teacher models provide pseudo-labels: an acoustic teacher (EnCodec model) that discretizes audio into tokens from 8 residual vector quantization codebooks, and a musical teacher based on constant-Q transform spectrogram reconstruction. MERT-v1-95M follows the HuBERT architecture with a CNN-based feature extractor and a 12-layer Transformer encoder producing 768-dimensional contextual embeddings.

The proposed two-stage CPT strategy addresses the stability gap phenomenon observed during continual pre-training. Stage 1 is a stabilization phase where only the CNN-based feature extractor and codeword embedding layer are updated while the Transformer encoder remains frozen, training on a smaller data subset with 20% Western music replay. Stage 2 is full adaptation where the Transformer encoder is unfrozen and training continues on the full dataset. Learning rate re-warming and re-decaying are applied in both stages, with Stage 1 using a moderately aggressive schedule and Stage 2 using a less aggressive schedule to balance plasticity and stability.

For task arithmetic, task vectors are computed as the element-wise difference between single-culture continually pre-trained models and the MERT-v1 model. A unified model is constructed by summing task vectors with corresponding scaling factors. The best task arithmetic variant was obtained using a scaling factor of λ = 0.2.

Evaluation results show that CultureMERT consistently outperforms the original MERT-v1 model across all non-Western tasks and evaluation metrics, achieving an average improvement of 4.9%. It also surpasses single-culture adapted models on average, suggesting that incorporating culturally diverse data during CPT benefits all non-Western traditions. CultureMERT achieves this with minimal forgetting on Western benchmarks (0.05% average drop across ROC-AUC and AP). Single-culture adapted models tend to perform best on their respective in-domain tasks for well-resourced traditions, and even low-resource adaptation (LyraMERT trained on just 50 hours) leads to noticeable gains across other non-Western tasks. Task arithmetic performs comparably to CultureMERT on non-Western tasks and even surpasses it on Western benchmarks and Lyra.

Cross-cultural transfer analysis reveals asymmetries in transfer effectiveness. Strong transfer is observed between Turkish-makam and Carnatic music, aligning with their shared theoretical foundations as modal frameworks emphasizing microtonality and improvisation. The Carnatic-adapted model appears most consistently transferable among single-culture adaptations, achieving strong results not only within Indian classical traditions but also generalizing well to Turkish-makam and Lyra.

Token-level culture similarity analysis using Jensen-Shannon divergence and cosine distance between token distributions extracted from the EnCodec model reveals strong token-level similarity among non-Western traditions, particularly between Hindustani and Carnatic music. Western datasets are highly similar to each other but notably dissimilar from non-Western traditions. Greek traditional music aligns more closely with non-Western traditions than Western ones. These findings correlate with cross-cultural transfer results, suggesting token-level similarity metrics can serve as predictors of positive cross-cultural transfer.

The paper acknowledges limitations including the frozen EnCodec tokenizer potentially being suboptimal for encoding culturally diverse musical languages, and suggests future directions including scaling to additional musical cultures, exploring alternative architectures, extending evaluation beyond sequence-level classification tasks, conducting fine-grained ablation studies, and investigating whether the two-stage CPT strategy remains necessary under less constrained computational budgets.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:

  1. Two-Stage Continual Pre-Training (CPT) Strategy
  • Implement learning rate re-warming and re-decaying in two distinct phases:

  • Stage 1: Freeze Transformer encoder, train only CNN feature extractor and codeword embeddings on a small subset (e.g., 100h) with 20% Western replay data, using a moderate warm-up (10%) and cosine decay.

  • Stage 2: Unfreeze all parameters, train on the full 650h multi-cultural dataset with a gentler schedule (1% warm-up, lower max LR).

  • This stabilizes training under limited compute (single 12GB GPU) and mitigates catastrophic forgetting.

  1. Task Arithmetic for Model Merging
  • Compute task vectors as the difference between single-culture adapted models (e.g., MakamMERT, CarnaticMERT) and the base MERT-v1-95M.

  • Merge these vectors with a shared scaling factor λ=0.2 (optimal per Figure 4) to create a unified multi-cultural model without additional training.

  1. Token-Level Similarity-Guided Data Mixing
  • Use Jensen-Shannon Divergence and cosine distance between EnCodec token distributions (as in Figure 3) to predict positive cross-cultural transfer.

  • Adjust pre-training data mixtures or task arithmetic weights based on these similarity scores, prioritizing cultures with higher token overlap (e.g., Hindustani-Carnatic).

  1. Probing-Based Evaluation Protocol
  • Freeze the adapted model and train only a shallow MLP (single 512-dim hidden layer) for auto-tagging, using 30-second sliding windows with averaged predictions.

  • This ensures fair, reproducible evaluation across diverse music traditions.

  • Cross-Cultural Music Understanding:

  • Achieve 4.9% average improvement in ROC-AUC and AP over MERT-v1 on non-Western auto-tagging tasks (Turkish-makam, Hindustani, Carnatic, Lyra), surpassing prior SOTA on all four.

  • Retain Western benchmark performance within 0.05% (minimal forgetting), so it works equally well on pop, rock, and classical.

  • Low-Resource Cultural Adaptation:

  • Adapt to a new musical tradition with as little as 50 hours of data (e.g., Lyra) and still improve cross-cultural generalization, as demonstrated by LyraMERT.

  • Training-Free Multi-Cultural Merging:

  • Combine multiple single-culture adapted models (e.g., MakamMERT + CarnaticMERT) into a unified model via task arithmetic, matching multi-cultural CPT performance without retraining—useful when compute is scarce or data cannot be shared.

  • Transfer Prediction and Data Curation:

  • Predict which cultures will benefit from shared pre-training using token similarity metrics, enabling smarter data selection for future adaptations (e.g., identify that Turkish-makam and Carnatic transfer well due to modal frameworks).

  • Deployment-Ready for Regional Applications:

  • Enable region-specific music recommendation, cultural heritage preservation, and non-Western music tagging in production, with open-source models (CultureMERT-95M and CultureMERT-TA-95M) that run on standard hardware.

Abstract

Recent advances in music foundation models have improved audio representation learning, yet their effectiveness across diverse musical traditions remains limited. We introduce CultureMERT-95M, a multi-culturally adapted foundation model developed to enhance cross-cultural music representation learning and understanding. To achieve this, we propose a two-stage continual pre-training strategy that integrates learning rate re-warming and re-decaying, enabling stable adaptation even with limited computational resources. Training on a 650-hour multi-cultural data mix, comprising Greek, Turkish, and Indian music traditions, results in an average improvement of 4.9% in ROC-AUC and AP across diverse non-Western music auto-tagging tasks, surpassing prior state-of-the-art, with minimal forgetting on Western-centric benchmarks. We further investigate task arithmetic, an alternative approach to multi-cultural adaptation that merges single-culture adapted models in the weight space. Task arithmetic performs on par with our multi-culturally trained model on non-Western auto-tagging tasks and shows no regression on Western datasets. Cross-cultural evaluation reveals that single-culture models transfer with varying effectiveness across musical traditions, whereas the multi-culturally adapted model achieves the best overall performance. To support research on world music representation learning, we publicly release CultureMERT-95M and CultureMERT-TA-95M, fostering the development of more culturally aware music foundation models.

Related papers