Musical Agent Systems: MACAT and MACataRT

summary

Video file (mp4)

In short

The episode discusses a paper on 'Musical Agent Systems: MACAT and MACataRT' by Keon Ju M. Lee and Philippe Pasquier from Simon Fraser University. The hosts explain that these AI systems function as musical partners, trained on small, personal datasets to create music that sounds like the artist's style, rather than generic formulas. They conclude that this research democratizes AI music creation by making it accessible on standard laptops and emphasizes its role as a collaborative instrument for musicians.

Key concepts

MACAT
MACAT is a musical agent designed to lead through self-listening. It uses a self-organizing map and a factor oracle to shape its own output, learning the structure of music rather than just individual sounds.
MACataRT
MACataRT is a collaborative improviser built on IRCAM’s CataRT system. It focuses on matching audio features in real time, making it excellent for reactive improvisation when jamming with a human musician.
Corpus-based concatenative synthesis
This technique involves analyzing small audio clips as puzzle pieces and stitching them together in real time based on pitch, timbre, and loudness. The system selects these pieces to create the musical output during performance.
Small, personal datasets
The agents are trained on limited data from the artist's own recordings. This approach ensures the AI learns the artist’s unique style and avoids relying on large, generic music pools, promoting ethical and transparent generative music.

Terminology used across episodes

This episode discusses

The paper

Musical Agent Systems: MACAT and MACataRT · Read on arXiv

Keon Ju M. Lee, Philippe Pasquier

Simon Fraser University

Our research explores the development and application of musical agents, human-in-the-loop generative AI systems designed to support music performance and improvisation within co-creative spaces. We introduce MACAT and MACataRT, two distinct musical agent systems crafted to enhance interactive music-making between human musicians and AI. MACAT is optimized for agent-led performance, employing real-time synthesis and self-listening to shape its output autonomously, while MACataRT provides a flexible environment for collaborative improvisation through audio mosaicing and sequence-based learning. Both systems emphasize training on personalized, small datasets, fostering ethical and transparent AI engagement that respects artistic integrity. This research highlights how interactive, artist-centred generative AI can expand creative possibilities, empowering musicians to explore new forms of artistic expression in real-time, performance-driven and music improvisation contexts.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Musical Agent Systems: MACAT and MACataRT".

Jane: The paper was written by Keon Ju M. Lee and Philippe Pasquier from Simon Fraser University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the channel, everyone. Today we’re digging into a fresh arXiv paper called “Musical Agent Systems: MACAT and MACataRT,” and I gotta say, the title alone got me curious.

Jane: Oh absolutely, Tom. The names sound like something out of a sci-fi movie, but they’re actually two very real AI systems built to make music with humans. The authors, Keon Ju M. Lee and Philippe Pasquier from Simon Fraser University, are part of the Metacreation Lab, which focuses on creative AI.

Tom: So these aren’t just tools that spit out a melody on command. They’re more like musical partners, right?

Jane: Exactly. The paper describes them as “musical agents,” which means they listen, learn, and respond in real time during a performance. MACAT is designed to lead, using self-listening to shape its own output. MACataRT is more of a collaborative improviser, great for jamming with a human musician.

Tom: And the coolest part? They train on small, personal datasets. So if you’re a percussionist, you feed it your own recordings, and the agent learns your style, not some generic pop formula.

Jane: That’s a huge deal. It means the AI isn’t just borrowing from a giant pool of unknown music. It’s working with material the artist actually owns and chooses. That’s a much more ethical and transparent way to do generative music.

Tom: I love that. It’s like having a bandmate who’s studied your playing and can react to you on stage, rather than a jukebox that plays random tracks.

Jane: And that’s the big implication here. This isn’t about replacing musicians. It’s about giving them a new kind of instrument, one that can improvise with them. The paper even mentions a live performance where MACataRT helped a percussionist and guitarist co-create a piece that made it to the finals of the two thousand twenty-four AI Music Song Contest.

Tom: Wait, they actually performed with this thing on stage? That’s not just a lab demo.

Jane: No, it’s real. They played in Zürich. So this paper is showing us that AI can be a genuine creative partner, not just a background tool. And that’s what I’m excited to dig into further.

Tom: Okay, so we’ve got the names and the big idea. But how do these systems actually work under the hood? That’s what we need to figure out next.

Summary: Jane: So, Tom, we’ve got the title and the vibe. Now let’s get into what the paper actually says about how these systems function. The core idea is that both MACAT and MACataRT use something called corpus-based concatenative synthesis.

Tom: That sounds like a mouthful. Can you break that down for me?

Jane: Sure. Imagine you have a big box of tiny audio clips, like little puzzle pieces. The system analyzes each piece for things like pitch, timbre, and loudness. Then, when it’s time to play, it picks pieces that fit the current musical moment and stitches them together in real time. That’s the “concatenative” part.

Tom: So it’s like a DJ with a massive crate of samples, but the DJ is an AI that’s listening to the room.

Jane: Exactly. And the key difference between the two systems is how they decide which piece to play. MACAT uses a self-organizing map, which is a way to cluster similar sounds on a 2D grid. It also uses a factor oracle, which is a pattern-matching algorithm that remembers sequences of sounds it has played before.

Tom: So MACAT is learning the structure of the music, not just the sounds.

Jane: Right. It’s like it remembers a chord progression and can predict what comes next. MACataRT, on the other hand, is built on IRCAM’s CataRT system, which is more about matching audio features in real time. It’s great for reactive improvisation, where the human plays and the AI responds.

Tom: And the paper says MACataRT can also be proactive, using that same factor oracle to generate music on its own.

Jane: Yes. So you have one system that’s great at leading and another that’s great at following, but both are trained on small, personal datasets. The paper emphasizes that this small-data approach keeps the AI aligned with the artist’s unique style.

Tom: That’s a big deal for musicians who don’t want to sound like everyone else. It’s not about generating generic music; it’s about generating music that sounds like you.

Jane: And that’s what makes this research so impactful. It’s shifting the conversation from big, impersonal models to personalized, collaborative tools. The paper even mentions that they use a “research-creation” methodology, which means they evaluate the systems through actual performances, not just computational metrics.

Tom: So they’re testing it on stage, with real audiences. That’s a pretty high bar.

Jane: It is. And it worked. The piece “Echoes of Synthetic Forest” by the duo KeRa was a finalist in the AI Music Song Contest. That’s proof that this isn’t just a toy; it’s a serious musical instrument.

Tom: Okay, so we’ve got the mechanics and the real-world validation. But what are the actual improvements this paper is suggesting? What’s new here compared to what came before?

Improvements: Jane: Alright Tom, so we know the basics. Now let’s talk about what this paper adds to the field. The big improvement here is the focus on temporality and explainability.

Tom: Temporality meaning how the AI handles time and musical structure?

Jane: Exactly. The original CataRT system was great at matching sounds but didn’t have a strong sense of musical sequence. MACataRT fixes that by adding the factor oracle, which learns patterns over time. So instead of just reacting to the last note, it can anticipate a phrase.

Tom: That’s like the difference between a drummer who just follows your beat and one who knows the song structure and can fill in the right spots.

Jane: Perfect analogy. And MACAT improves on the older MASOM system by adding real-time sound synthesis and better visualization. So the artist can see which nodes are being played and tweak parameters like pitch and resampling on the fly.

Tom: So it’s not just a black box. You can actually see what the AI is thinking.

Jane: Right, and that ties into explainability. The paper proposes adding a module that records the history of past musical patterns. That way, you can trace why the AI made a certain choice, which is huge for building trust between the musician and the system.

Tom: And they’re also planning to use deep learning for longer sequences, right?

Jane: Yes. The current factor oracle is good, but it’s limited in how far back it can remember. They want to use neural networks to learn longer musical structures, like entire bars or phrases. That would make the AI feel even more like a bandmate who remembers the whole song.

Meng: Hey, Jane, can I jump in here? I’m curious about the practical side. You mentioned they train on small datasets. How small are we talking?

Jane: The paper mentions one example where the trained model had ninety-nine audio segments organized on a four by four self-organizing map. So we’re talking about a very manageable amount of data.

Meng: And they don’t need a GPU? That’s wild. The paper says they trained on just a MacBook CPU. That means a musician could do this at home without any special hardware.

Jane: Exactly. That’s a huge barrier removed. You don’t need a data center to make personalized AI music. You just need your own recordings and a laptop.

Lu: And that’s what makes this so exciting from a research perspective. It’s democratizing AI. You’re not relying on some massive, anonymous dataset. You’re using your own artistic voice to train the model. That’s a fundamentally different relationship between the creator and the tool.

Tom: So the improvements here are about making the AI more musical, more understandable, and more accessible. That’s a pretty solid trio.

Jane: And it’s all grounded in real performances. They’re not just theorizing; they’re showing it works on stage. That’s the kind of research that can actually change how musicians work.

Tom: Alright, so we’ve covered the mechanics and the improvements. Let’s wrap this up and think about what it all means for the future.

Conclusion: Tom: We’ve been talking about “Musical Agent Systems: MACAT and MACataRT” all episode, and I think we’ve only scratched the surface. Let’s pull it all together.

Jane: Absolutely. The paper introduces two musical agents that act as creative partners. MACAT leads with self-listening, and MACataRT collaborates through audio mosaicing. Both are trained on small, personal datasets, which keeps the music authentic and ethical.

Tom: And the big takeaway for me is that this isn’t about replacing human creativity. It’s about expanding it. The systems are designed to be instruments, not composers. The musician is still in control.

Jane: Right. And the paper’s future work is exciting. They want to add deep learning for longer musical memory and a feedback loop using reinforcement learning to make the agents even more adaptive in live settings.

Lu: I think the most impactful part is the ethical framework. By using small datasets, they avoid the copyright issues that plague big AI models. They’re respecting the artist’s voice and their intellectual property.

Meng: And from a practical standpoint, the fact that it runs on a regular laptop means this could actually be adopted by working musicians. It’s not a lab curiosity; it’s a tool.

Lalam: And culturally, this could change how we think about AI in the arts. Instead of a threat, it becomes a collaborator that helps artists explore new sonic territories while staying true to their own style. That’s a positive vision for human-AI co-creation.

Tom: Well said. So we’re saying goodbye to MACAT and MACataRT, but I have a feeling we’ll be hearing more about them. Thanks for joining us, and we’ll see you on the next paper.

More episodes

← Home