Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

summary

Video file (mp4)

The gist

The paper, "Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset," addresses the limitations in existing Audio-to-Score (A2S) research, which

In short

This episode discusses the 'Audio-to-Score Transcription' paper, which introduces the SheetSage-A2S dataset—61 hours of audio paired with kern score encodings. The method uses data augmentation and MuQ features to improve robustness against real-world variations. Results show a 4.98% Symbol Error Rate (SER) on classical music and a 20.92% SER on popular music, providing an accessible tool for AI musical comprehension.

Key concepts

SheetSage-A2S Dataset
This is a massive collection of sixty-one hours of audio clips from over six thousand unique songs. The dataset provides the ground truth for researchers by pairing raw audio with precise kern score encodings, allowing AI to learn from real-world musical texture.
Data Augmentation
The authors generate augmented variants by shifting pitch and stretching time for every training example. This technique makes the transcription model more robust, ensuring it can handle real-world variations in performance, such as slight differences in tempo or pitch.
MuQ (Pre-trained Feature Extractor)
MuQ replaces basic feature extractors to provide high-level understanding of music structure. It leverages learned musical intelligence from a large dataset, allowing the AI to understand concepts like pitch and harmony rather than just low-level audio features.
Symbol Error Rate (SER)
This metric measures the accuracy of the transcription system. The model achieved a 4.98% SER when testing classical music using the Quartets collection, and a 20.92% SER when applying SheetSage-A2S to popular music.

Terminology used across episodes

This episode discusses

The paper

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset · Read on arXiv

Haina Zhu, Yizhi Zhou, Hangting Chen, Jianwei Yu, Ziyang Ma, Rongzhi Gu, Yi Luo, Wei Tan, Xie Chen

IEEE Transactions on Audio, Speech and Language Processing

Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with**kern score encodings for 9,468 clips originating from 6,066 unique songs, the first of its kind to facilitate A2S research for popular music. Additionally, we improve on existing A2S approaches by using data augmentation and MuQ, a pretrained feature-extraction model for music audio, to enhance generalisation abilities and extract meaningful audio features. Results show that the proposed A2S model achieves 4.98% symbol error rate (SER) on the Quartets collection for classical music, which significantly outperforms the 15.3% SER from the existing state-of-the-art. Additionally, our model achieves 20.92% SER on the SheetSage-A2S dataset for popular music, serving as a strong benchmark for future research. The dataset, model, and code are made publicly available at: https://github.com/Multimodal-Music-Research-Lab/SheetSage2Kern model.

DOI: 10.1145/3767308.3835653

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset".

Jane: The paper was written by Haina Zhu, Yizhi Zhou, Hangting Chen, Jianwei Yu, Ziyang Ma et al. from IEEE Transactions on Audio, Speech and Language Processing and IEEE.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary and Core Methodology: Tom: So, the paper presents the SheetSage-A2S Dataset, which is a huge collection of sixty-one hours of audio clips from over six thousand unique songs. That's a massive amount of data for this specific task.

Jane: And it’s not just raw audio; they are pairing that audio with precise kern score encodings, which is the format used to represent the sheet music. This provides the ground truth that researchers need to train their models.

Lu: The dataset contains nine thousand four hundred sixty-eight clips in total, and I think that number speaks volumes about how much real-world data they managed to collect compared to previous datasets like Quartets or Chorales.

Meng: Using kern as the format makes sense because it’ is compact and effective for representing those lead sheets without the complexity of a full MusicXML file, which simplifies the training targets.

Lalam: It feels like this dataset is providing a foundation that allows us to understand how AI can learn from the texture of real music instead of just relying on synthesized examples.

Tom: The authors also highlight two major improvements over existing A2S approaches in this summary, which are using data augmentation and a pre-trained feature extractor called MuQ.

Jane: They’re suggesting that by feeding the model augmented data—like time stretching or pitch shifting—we can make the transcription much more robust to handle real-world variations in performance.

Lu: I think that's a critical point, because if we are only training on perfect recordings, the model will fail when trying to transcribe something slightly off-key or with a different tempo.

Meng: And replacing the original approach with MuQ seems like an attempt to use high-level understanding of music rather than just low-level audio features, which is a practical way to boost performance.

Lalam: When we start seeing AI models trained on such a vast and diverse dataset, we' are moving toward a level of musical comprehension that feels almost intuitive.

Technical Improvements and Generalization: Tom: The paper details three specific modifications to their A2S model, which are designed to boost generalization significantly. They’re making changes to the architecture itself.

Jane: One key change is expanding the feedforward dimension of the transformer decoder from two hundred fifty-six up to one thousand twenty-four and switching to a Pre-Norm architecture for stability. This allows the model more capacity to learn complex relationships in music.

Lu: I think switching to Pre-Norm is essential because it addresses training instability, which is a common problem when we scale up deep learning models for tasks as complicated as music transcription.

Meng: And the move from the vanilla CNN encoder to MuQ really speaks to practical impact; you’re essentially swapping a basic feature extractor for a sophisticated model that understands music structure already knows how to see pitch and harmony.

Lalam: It's fascinating how they are leveraging a pre-trained model, which means the AI isn't just learning note timings; it’s leveraging learned musical intelligence from the entire Million Song Dataset.

Tom: This architectural upgrade is coupled with data augmentation, where they generate six augmented variants for every single training example by shifting pitch and stretching time.

Jane: That helps the model become more resilient to real-world recording differences, so it doesn's not just memorizing specific beats but actually understanding the musical content.

Lu: It’s a sophisticated way of telling you, "Don't worry about whether this note was played slightly fast or slow me down, learn the concept."

Meng: I wonder how much of that complexity is necessary for a real-world implementation; does this need to run on specialized hardware to handle all those twelve hypotheses and multiple augmentations?

Lalam: The whole system is designed so that we can finally move past the limitations of one single, specific sound and embracing the diverse ways human performance can express musical ideas.

Results, Implications, and Benchmark: Tom: Let's look at the results; they found that with their proposed model, they achieved a Symbol Error Rate of four point nine eight percent on classical music for the Quartets collection. That's a massive improvement over the previous state-of-the-art.

Jane: And then, when looking at popular music with SheetSage-A2S, they achieved an SER of twenty point nine two percent. While that might sound higher than the classical results, it sets a strong benchmark for what is possible in a genre that was previously ignored.

Lu: The fact that they are able to get such low error rates on the Quartets collection validates their approach, but the twenty point nine two percent on SheetSage-A2S shows they aren’t just solving one problem at all, Meng.

Meng: From an engineering perspective, that twenty point nine two percent is a very practical figure; it tells us that for popular music—where the structure is less rigid than classical music—this A2S system provides a usable, high-quality transcription tool for real applications.

Lalam: The implications are huge because this means the AI isn't just an academic curiosity; it’s becoming a practical tool that allows people to interact with and study our cultural output in a new ways.

Tom: It’s not just about the accuracy either; they are providing the code and data for SheetSage-A2S, making it open and accessible for future researchers to build on.

Jane: That means the this work is really accelerating how we learn to do A2S, letting other people start experimenting right with this new data.

Lu: I think the authors are setting a standard that will force subsequent models to push past current limitations, defining a clear path forward for their research field.

Meng: A clear path indeed; it gives developers a concrete target and also shows them what level of performance is actually achievable in this domain.

Lalam: The world is going to be much richer when the technology can accurately transcribe not just Bach or Mozart, but also the popular music that defines our lives today.

Conclusion and Wrap-up: Tom: We’ve spent quite a bit of time discussing "Audio-to-Score Transcription using Pre-Trained Features, Data Augmentation, and the New SheetSage-A2S Dataset," and it's clear that this is a foundational piece of work.

Jane: The core of our discussion was how successfully tackling the data scarcity in popular music by presenting sixty-one hours of audio has opened up a new frontier for AI to understand musical notation.

Lu: I see this as a major milestone because the authors have not only solved a technical problem but also demonstrated how to handle real-world, user-generated content with minimal preprocessing.

Meng: My takeaway is that this system provides high-quality outputs—a twenty point nine two percent SER on popular music—which is exactly what we need for practical applications in transcription and archival software.

Lalam: The ultimate impact, as I see it, is that the technology will enable a deeper cross-cultural appreciation of music by providing accurate scores for all forms of musical expression.

Tom: Before we wrap up, let's hear one final thought from each of you on this incredible achievement.

Jane: This is a huge win for making music accessible to everyone who appreciates it, not just the specialized classical audience.

Lu: I believe this will be the starting point for many more sophisticated models that incorporate learning from real-world musical performance characteristics.

Meng: I’m confident that seeing how MuQ and data augmentation perform so well means this is a scalable architecture for real-time transcription too, in my opinion.

Lalam: The world is poised to become a richer place as we can finally bridge the gap between the sound of popular music and its symbolic representation.

Tom: Thank you all so much for joining us on this exciting journey into "Audio-to-Score Transcription using Pre-Trained Features, Data Augmentation, and the New SheetSage-A2S Dataset." We’ll see you next time!

More episodes

← Home