Learning and composing of classical music using restricted Boltzmann machines

summary

Video file (mp4)

The gist

This study investigates how machine learning models acquire the ability to compose music and how musical information is internally represented within such models.

In short

The episode discusses a paper by Kobayashi and Watanabe on using restricted Boltzmann machines (RBMs) to learn and compose classical music from piano roll images of Bach and Mozart. The hosts explore how the simple model can generate two-measure coherent pieces but fails at longer structures. They conclude that while the model learns rhythm, its internal representations do not map onto human music theory.

Key concepts

Restricted Boltzmann Machine (RBM)
A simple type of neural network that learns patterns from data. In this study, it was used to analyze piano roll images of classical music to see if a minimal model could learn musical structure.
Piano Roll Images
Black-and-white images created by converting musical scores. Time runs horizontally and pitch runs vertically, allowing the model to look at music like a picture.
Interpretability
The authors prioritized understanding what the model learns inside its structure over just achieving high generative performance. The hosts discuss how this lack of interpretability relates to building trust in AI.

Terminology used across episodes

This episode discusses

The paper

Learning and composing of classical music using restricted Boltzmann machines · Read on arXiv

Mutsumi Kobayashi, Hiroshi Watanabe

Keio University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Learning and composing of classical music using restricted Boltzmann machines".

Jane: The paper was written by Mutsumi Kobayashi and Hiroshi Watanabe from Keio University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making the rounds on arXiv, and it’s called “Learning and composing of classical music using restricted Boltzmann machines.” Jane, I have to say, the title alone brings me back to my old machine learning classes.

Jane: It does, Tom, and that’s exactly why I love it. So, a restricted Boltzmann machine, or RBM for short, is basically a simple type of neural network that learns patterns from data. The authors here, Mutsumi Kobayashi and Hiroshi Watanabe from Keio University, they took piano roll images of Bach’s music and fed them into this RBM.

Tom: Right, and a piano roll is that old-school player piano format, right? The paper converts musical scores into these black-and-white images where time runs horizontally and pitch runs vertically. So the model is essentially learning to look at music like it’s a picture.

Jane: Exactly. And the big deal here is that they wanted to see if this really simple model could compose music, and more importantly, what it actually learns inside. They’re not trying to build the next big AI composer, they want to open up the black box.

Tom: And that’s what gets me excited. They’re asking, “Can a minimal model generate something musical, and can we peek inside to see how it’s doing it?” That’s a really refreshing approach compared to the giant deep learning models we usually hear about.

Jane: Totally. And the paper’s authors are pretty upfront that they’re prioritizing interpretability over sheer generative performance. They want to use the model as a tool to understand musical structure, not just to churn out pieces.

Tom: So, for our listeners, this is a paper about using a very transparent, simple AI to compose classical music, and then dissecting that AI to see what musical concepts it picked up on its own. It’s like giving a kid a box of LEGOs and then asking them to explain their building rules afterwards.

Jane: That’s a great way to put it. And the implications are pretty big for how we think about AI in creative fields. If a simple model can learn musical structure, what does that say about what’s necessary for creativity? We’ll get into the actual results in a moment.

Tom: Stick around, because we’re about to see if this little RBM can actually hold a tune.

Paper discussion segment 2: Jane: Alright, we’re back, and we’re still on “Learning and composing of classical music using restricted Boltzmann machines.” So, Tom, we set the stage, but what did the authors actually find?

Tom: Well, Jane, the first thing they did was test if the model could reconstruct piano rolls. They fed it a Bach piece it had seen before, and it rebuilt it perfectly. Then they fed it a Mozart piece it had never seen, and it still reconstructed it accurately. That’s a good sign that it learned general musical features, not just memorized the training data.

Jane: And here’s the kicker—they fed it MNIST digit images, you know, the handwritten numbers. The model completely failed to reconstruct those. It just produced noise. So the RBM definitely knows what a piano roll looks like, and it knows what it doesn’t look like.

Tom: That’s such a clean experiment. They even computed the energy of the model for different inputs. Piano rolls had really low energy, like negative three thousand six hundred while the digit images and noise had much higher, sometimes positive, energy. It’s like the model has a built-in “musicality” meter.

Jane: Right, and lower energy in a Boltzmann machine means the model considers that input to be more probable. So it’s essentially saying, “Yes, this looks like the music I learned, and that digit image looks like nonsense to me.”

Tom: Then comes the fun part—they actually had it compose. They developed an algorithm to generate new piano rolls. The two-measure pieces it created were surprisingly coherent. The paper notes that the generated segment was in B minor, contained a proper E minor chord, and even had a nice stepwise melody.

Jane: And that’s the part that made me sit up. They weren’t just generating random notes; the model was producing something that follows musical rules, like harmony and key. But then they tried to extend it to eight measures, and that’s where things fell apart.

Tom: Yeah, the longer pieces started out okay, in F major, but after a few measures the pitch content got messy and the musical coherence just dissolved. So the model can handle two measures, which is what it was trained on, but it can’t really hold a longer narrative together.

Jane: That’s a really honest result. It shows the limits of the model, but it also tells us something important. The RBM learned local musical grammar, like chords and short melodic phrases, but it didn’t learn long-term structure. It’s like it knows how to write a good sentence but not a good paragraph.

Tom: And that’s a huge insight for anyone building music AI. It suggests that long-range structure needs a different mechanism, maybe something with memory. But for a simple model, getting the local grammar right is already pretty impressive.

Jane: Exactly. And this sets us up perfectly for the next part, where we talk about what the authors found when they looked inside the model’s brain, so to speak.

Paper discussion segment 3: Tom: Welcome back. We’re still on “Learning and composing of classical music using restricted Boltzmann machines,” and now we get to the part I’ve been waiting for—the internal representations. Jane, what did they see when they poked around inside?

Jane: So, they did something really clever. They activated individual hidden units in the model, one at a time, and looked at what kind of visible pattern that unit was responsible for. It’s like asking each neuron, “What are you looking for in the music?”

Tom: And what did they find? I’m guessing it wasn’t a perfect “this is a C major chord” detector.

Jane: You’d be right. What they found were a lot of local temporal patterns, mostly related to note duration. The hidden units had learned to recognize things like sixteenth-note rhythms. So the model was really good at understanding rhythm, which makes sense because that’s a very visual, structural feature of a piano roll.

Tom: But they didn’t find anything that looked like a melodic phrase or a chord shape, right? The paper says those were barely observed. So the model’s internal language isn’t the same as our music theory language.

Jane: Exactly. And that’s a really profound finding. The model is learning the statistical structure of the data, but it’s not learning it in a way that maps neatly onto human concepts like “dominant seventh chord” or “cadence.” It’s a data-driven representation that’s fundamentally different from how we think about music.

Tom: And that has huge implications for explainable AI. We want to trust these models, but if their internal logic is alien to us, how do we build that trust? The authors are essentially saying, “Hey, we looked, and it’s not interpretable in the way we hoped.”

Jane: Right. But they also suggest this isn’t necessarily a bad thing. The latent space of the RBM could be used as a new analytical tool. It’s a way to look at music that isn’t bound by traditional theory. Maybe it can find patterns we haven’t noticed.

Tom: That’s a really creative spin on it. Instead of forcing the model to explain itself in our terms, we learn to read its terms. And the paper also points out that this lack of interpretability might be because the RBM’s representations are complex mixtures of features. They mention that classification RBMs might produce more distinguishable units.

Jane: So there’s a clear path forward. They’re not just saying, “It’s a black box, give up.” They’re saying, “Here’s what we found, and here’s how we might build a model that’s more interpretable.” That’s the kind of constructive criticism we need in this field.

Tom: And it makes me wonder about the bigger picture. If a simple model learns rhythm but not harmony in a human-readable way, what does that say about how we teach music? Maybe our theoretical framework is just one way to slice the pie.

Jane: That’s a great question, Tom, and I think we should bring in the rest of the team to chew on that. But first, let’s wrap up our thoughts on this paper.

Conclusion: Jane: Alright, we’re wrapping up our discussion on “Learning and composing of classical music using restricted Boltzmann machines.” Tom, give us the final takeaway.

Tom: Sure, Jane. The paper showed that a simple RBM can learn to compose two-measure pieces of classical music that are musically coherent, at least locally. It can even distinguish between musical and non-musical images. But when they tried to generate longer pieces, the coherence fell apart.

Jane: And the most important finding was that the model’s internal representations don’t match human music theory. It learned rhythm patterns, but not recognizable chords or melodies. That tells us that a minimal generative model can capture statistical regularities, but its way of organizing that information is fundamentally different from ours.

Tom: Right. And that’s a big deal for the field of explainable AI in creative tasks. It’s a reminder that just because a model can produce something that sounds good, doesn’t mean we can easily understand how it’s doing it. The authors did a great job of being honest about that limitation.

Jane: And they also gave us a roadmap for the future. They suggested looking at different composers, analyzing the weight matrices, and trying different architectures like classification RBMs. So this isn’t a dead end, it’s a starting point for building more interpretable music AI.

Tom: Absolutely. And on a personal note, I love that they made their code available on GitHub. That’s how science moves forward—by letting others build on your work.

Jane: Couldn’t agree more. So, to the authors, Kobayashi and Watanabe, thank you for this thoughtful paper. To our listeners, we hope you enjoyed this deep dive. We’re going to say goodbye to this paper and get ready to look at the next one.

Tom: See you all next time, and keep listening to the music of the data.

Jane: Take care, everyone.

More episodes

← Home