Musical Agent Systems: MACAT and MACataRT
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Musical Agent Systems: MACAT and MACataRT".
Jane: The paper was written by Keon Ju M. Lee and Philippe Pasquier from Simon Fraser University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the channel, everyone. Today we’re digging into a fresh arXiv paper called “Musical Agent Systems: MACAT and MACataRT,” and I gotta say, the title alone got me curious.
Jane: Oh absolutely, Tom. The names sound like something out of a sci-fi movie, but they’re actually two very real AI systems built to make music with humans. The authors, Keon Ju M. Lee and Philippe Pasquier from Simon Fraser University, are part of the Metacreation Lab, which focuses on creative AI.
Tom: So these aren’t just tools that spit out a melody on command. They’re more like musical partners, right?
Jane: Exactly. The paper describes them as “musical agents,” which means they listen, learn, and respond in real time during a performance. MACAT is designed to lead, using self-listening to shape its own output. MACataRT is more of a collaborative improviser, great for jamming with a human musician.
Tom: And the coolest part? They train on small, personal datasets. So if you’re a percussionist, you feed it your own recordings, and the agent learns your style, not some generic pop formula.
Jane: That’s a huge deal. It means the AI isn’t just borrowing from a giant pool of unknown music. It’s working with material the artist actually owns and chooses. That’s a much more ethical and transparent way to do generative music.
Tom: I love that. It’s like having a bandmate who’s studied your playing and can react to you on stage, rather than a jukebox that plays random tracks.
Jane: And that’s the big implication here. This isn’t about replacing musicians. It’s about giving them a new kind of instrument, one that can improvise with them. The paper even mentions a live performance where MACataRT helped a percussionist and guitarist co-create a piece that made it to the finals of the two thousand twenty-four AI Music Song Contest.
Tom: Wait, they actually performed with this thing on stage? That’s not just a lab demo.
Jane: No, it’s real. They played in Zürich. So this paper is showing us that AI can be a genuine creative partner, not just a background tool. And that’s what I’m excited to dig into further.
Tom: Okay, so we’ve got the names and the big idea. But how do these systems actually work under the hood? That’s what we need to figure out next.
Summary: Jane: So, Tom, we’ve got the title and the vibe. Now let’s get into what the paper actually says about how these systems function. The core idea is that both MACAT and MACataRT use something called corpus-based concatenative synthesis.
Tom: That sounds like a mouthful. Can you break that down for me?
Jane: Sure. Imagine you have a big box of tiny audio clips, like little puzzle pieces. The system analyzes each piece for things like pitch, timbre, and loudness. Then, when it’s time to play, it picks pieces that fit the current musical moment and stitches them together in real time. That’s the “concatenative” part.
Tom: So it’s like a DJ with a massive crate of samples, but the DJ is an AI that’s listening to the room.
Jane: Exactly. And the key difference between the two systems is how they decide which piece to play. MACAT uses a self-organizing map, which is a way to cluster similar sounds on a 2D grid. It also uses a factor oracle, which is a pattern-matching algorithm that remembers sequences of sounds it has played before.
Tom: So MACAT is learning the structure of the music, not just the sounds.
Jane: Right. It’s like it remembers a chord progression and can predict what comes next. MACataRT, on the other hand, is built on IRCAM’s CataRT system, which is more about matching audio features in real time. It’s great for reactive improvisation, where the human plays and the AI responds.
Tom: And the paper says MACataRT can also be proactive, using that same factor oracle to generate music on its own.
Jane: Yes. So you have one system that’s great at leading and another that’s great at following, but both are trained on small, personal datasets. The paper emphasizes that this small-data approach keeps the AI aligned with the artist’s unique style.
Tom: That’s a big deal for musicians who don’t want to sound like everyone else. It’s not about generating generic music; it’s about generating music that sounds like you.
Jane: And that’s what makes this research so impactful. It’s shifting the conversation from big, impersonal models to personalized, collaborative tools. The paper even mentions that they use a “research-creation” methodology, which means they evaluate the systems through actual performances, not just computational metrics.
Tom: So they’re testing it on stage, with real audiences. That’s a pretty high bar.
Jane: It is. And it worked. The piece “Echoes of Synthetic Forest” by the duo KeRa was a finalist in the AI Music Song Contest. That’s proof that this isn’t just a toy; it’s a serious musical instrument.
Tom: Okay, so we’ve got the mechanics and the real-world validation. But what are the actual improvements this paper is suggesting? What’s new here compared to what came before?
Improvements: Jane: Alright Tom, so we know the basics. Now let’s talk about what this paper adds to the field. The big improvement here is the focus on temporality and explainability.
Tom: Temporality meaning how the AI handles time and musical structure?
Jane: Exactly. The original CataRT system was great at matching sounds but didn’t have a strong sense of musical sequence. MACataRT fixes that by adding the factor oracle, which learns patterns over time. So instead of just reacting to the last note, it can anticipate a phrase.
Tom: That’s like the difference between a drummer who just follows your beat and one who knows the song structure and can fill in the right spots.
Jane: Perfect analogy. And MACAT improves on the older MASOM system by adding real-time sound synthesis and better visualization. So the artist can see which nodes are being played and tweak parameters like pitch and resampling on the fly.
Tom: So it’s not just a black box. You can actually see what the AI is thinking.
Jane: Right, and that ties into explainability. The paper proposes adding a module that records the history of past musical patterns. That way, you can trace why the AI made a certain choice, which is huge for building trust between the musician and the system.
Tom: And they’re also planning to use deep learning for longer sequences, right?
Jane: Yes. The current factor oracle is good, but it’s limited in how far back it can remember. They want to use neural networks to learn longer musical structures, like entire bars or phrases. That would make the AI feel even more like a bandmate who remembers the whole song.
Meng: Hey, Jane, can I jump in here? I’m curious about the practical side. You mentioned they train on small datasets. How small are we talking?
Jane: The paper mentions one example where the trained model had ninety-nine audio segments organized on a four by four self-organizing map. So we’re talking about a very manageable amount of data.
Meng: And they don’t need a GPU? That’s wild. The paper says they trained on just a MacBook CPU. That means a musician could do this at home without any special hardware.
Jane: Exactly. That’s a huge barrier removed. You don’t need a data center to make personalized AI music. You just need your own recordings and a laptop.
Lu: And that’s what makes this so exciting from a research perspective. It’s democratizing AI. You’re not relying on some massive, anonymous dataset. You’re using your own artistic voice to train the model. That’s a fundamentally different relationship between the creator and the tool.
Tom: So the improvements here are about making the AI more musical, more understandable, and more accessible. That’s a pretty solid trio.
Jane: And it’s all grounded in real performances. They’re not just theorizing; they’re showing it works on stage. That’s the kind of research that can actually change how musicians work.
Tom: Alright, so we’ve covered the mechanics and the improvements. Let’s wrap this up and think about what it all means for the future.
Conclusion: Tom: We’ve been talking about “Musical Agent Systems: MACAT and MACataRT” all episode, and I think we’ve only scratched the surface. Let’s pull it all together.
Jane: Absolutely. The paper introduces two musical agents that act as creative partners. MACAT leads with self-listening, and MACataRT collaborates through audio mosaicing. Both are trained on small, personal datasets, which keeps the music authentic and ethical.
Tom: And the big takeaway for me is that this isn’t about replacing human creativity. It’s about expanding it. The systems are designed to be instruments, not composers. The musician is still in control.
Jane: Right. And the paper’s future work is exciting. They want to add deep learning for longer musical memory and a feedback loop using reinforcement learning to make the agents even more adaptive in live settings.
Lu: I think the most impactful part is the ethical framework. By using small datasets, they avoid the copyright issues that plague big AI models. They’re respecting the artist’s voice and their intellectual property.
Meng: And from a practical standpoint, the fact that it runs on a regular laptop means this could actually be adopted by working musicians. It’s not a lab curiosity; it’s a tool.
Lalam: And culturally, this could change how we think about AI in the arts. Instead of a threat, it becomes a collaborator that helps artists explore new sonic territories while staying true to their own style. That’s a positive vision for human-AI co-creation.
Tom: Well said. So we’re saying goodbye to MACAT and MACataRT, but I have a feeling we’ll be hearing more about them. Thanks for joining us, and we’ll see you on the next paper.
Keon Ju M. Lee, Philippe Pasquier
Simon Fraser University
cs.MA, cs.AI, cs.HC, cs.SD, eess.AS
Submitted: 2026-08-14
Updated: 2026-08-17
Comments: In Proceedings of the Creativity and Generative AI NIPS (Neural Information Processing Systems) Workshop 2024
Code: https://github.com/Metacreation-Lab/Musical-Agent-Systems
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 46/100
Key concepts
- MACAT
- MACAT is a musical agent designed to lead through self-listening. It uses a self-organizing map and a factor oracle to shape its own output, learning the structure of music rather than just individual sounds.
- MACataRT
- MACataRT is a collaborative improviser built on IRCAM’s CataRT system. It focuses on matching audio features in real time, making it excellent for reactive improvisation when jamming with a human musician.
- Corpus-based concatenative synthesis
- This technique involves analyzing small audio clips as puzzle pieces and stitching them together in real time based on pitch, timbre, and loudness. The system selects these pieces to create the musical output during performance.
- Small, personal datasets
- The agents are trained on limited data from the artist's own recordings. This approach ensures the AI learns the artist’s unique style and avoids relying on large, generic music pools, promoting ethical and transparent generative music.
Terminology
Summary
Summary
This paper introduces and details two distinct musical agent systems developed by the Metacreation Lab: MACAT and MACataRT. These systems are described as human-in-the-loop generative AI systems designed to support music performance and improvisation within co-creative spaces.
The research is motivated by the goal of developing interactive, artist-in-the-loop generative AI systems designed for musicians and sound artists
to enable exploration of novel creative possibilities in real-time and interactive environments.
The systems are based on principles of using small data in music
and model crafting,
which allows artists to train models on curated, personalized datasets. This approach contrasts with large-scale models, as it enables our musical agent systems to closely align with the specific musical nuances and stylistic preferences of collaborating musicians.
The paper emphasizes that this focused approach enhances the agent’s capacity to function as a genuine creative collaborator, fostering stylistic coherence and adaptability.
MACAT is described as an enhanced version of the MASOM agent
and is optimized for agent-led performance, employing real-time synthesis and self-listening to shape its output autonomously.
Its workflow integrates concatenative sound synthesis,
a Self-Organizing Map (SOM) for clustering audio segments by timbral similarity, and a Factor Oracle (FO) for real-time pattern recognition. During training, MACAT performs offline machine listening to analyze the original audio data, initialize SOM nodes, and identify the Best Matching Unit (BMU) for each input vector based on Euclidean distance.
In real-time generation, it uses a congruence parameter to regulate FO’s forward and backward transitions, ensuring responsive and context-sensitive generative output.
MACAT was showcased at the MusicAcoustica Festival in Hangzhou, China, by the artist collective K-Phi-A.
MACataRT is an enhanced version of the CataRT system, incorporating a temporal model based on the factor oracle and offering a more intuitive interface for sound synthesis and resampling.
It provides a flexible environment for collaborative improvisation through audio mosaicing and sequence-based learning.
The system supports two modes: reactive improvisation,
which responds to live inputs based on machine listening without using the factor oracle, and proactive improvisation,
where the agent learns sequences of audio segment indices during offline training and uses the factor oracle to generate music. MACataRT's practical effectiveness was demonstrated in a live performance by the duo KeRa, whose piece Echoes of Synthetic Forest
was selected as a Top 10 finalist in the 2024 AI Music Song Contest.
The paper details the workflows of both systems, contrasting them with their predecessors (MASOM and CataRT). It highlights that while CataRT lacks a temporal model to manage time-based musical structures,
MACataRT integrates the factor oracle to automate the generation process. The research employs a research-creation methodology,
which combines scholarly inquiry with creative experimentation, focusing on co-creation between humans and AI systems.
The authors argue this approach is essential because quantitative analysis alone cannot fully capture
the interpretive, spontaneous, and affective qualities that emerge through real-time collaboration.
The paper also extensively discusses ethical implications in Appendix A. The systems' use of small, personalized datasets of music recordings created by specific musicians
is presented as an advantage over large-scale models, promoting ethical use of AI in generative music by respecting the intellectual property rights and artistic contributions of other composers.
This approach enhances transparency and accountability of the creative process
and allows for straightforward tracking of how specific pieces of music contribute to and shape the generated content.
Additionally, training on smaller datasets is noted as more environmentally friendly
as it does not require external GPUs for training
and reduces energy consumption.
For future work, the authors aim to enhance both the temporality and explainability of our musical agent systems
by integrating deep learning architectures that enable agents to learn and retain longer sequences of musical patterns
and incorporating a module that records the history of past musical patterns.
They also plan to advance the machine listening module and potentially introduce a feedback loop using reinforcement learning to further enhance the adaptability of musical agents in real-time performance contexts.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems, along with the resulting capabilities:
Improvement: Replace the default large-scale dataset training with a two-stage pipeline: (a) a corpus curation stage where the artist selects and segments their own audio recordings (e.g., 99 segments as in the paper), and (b) a training stage using only that corpus, with no external GPUs (runs on CPU). This uses the paper’s model crafting
principle.
Improved AI capability: The AI can be trained on a specific musician’s personal recordings (e.g., a percussionist’s 5-minute improvisation) to generate output that matches their timbral and rhythmic nuances, without needing cloud computing or large datasets. This enables real-time personalization on a laptop, reducing carbon footprint and ensuring the AI does not mimic unlicensed artists.
Improvement: Integrate a Factor Oracle (suffix automaton) into the AI’s generation loop, as done in MACataRT. This model learns sequences of audio segment indices during offline training and uses forward/backward jumps controlled by a congruence
parameter (0.0–1.0) to balance repetition vs. novelty.
Improvement: Add a real-time audio feature extraction module (as in MASOM/MACAT) that computes 13 MFCCs, loudness, spectral flatness (4 bands), perceptual spectral decrease, and perceptual tristimulus, plus valence/arousal from the Circumplex model. This module resets statistics on each new sample and feeds a 31-dimensional vector to the generation module.
Improvement: Implement a switchable mode, as in MACataRT: (a) reactive mode uses K-Nearest Neighbors (KNN) to match live input features (e.g., centroid, periodicity) and selects corpus segments accordingly, without temporal modeling; (b) proactive mode uses the trained FO to generate sequences, with the option to blend in reactive responses.
Improvement: Add a 2D self-organizing map (SOM) visualization that highlights the current node (high luminance) and the previous node (low luminance), as in MACAT. Also, display the currently played node number, sample duration, and the artist/song name from the corpus.
Improvement: Add independent controls for resampling (simultaneously changes playback speed and pitch) and pitch shift (in cents, independent of speed), as in MACAT’s interface. Also, include reverse playback and fade-in/out (attack/release) to prevent clipping.
Improvement: Expose a congruence parameter (0.0–1.0) that directly controls the probability of forward vs. backward jumps in the Factor Oracle. At 1.0, the AI repeats the same node; at 0.0, it jumps to unrelated nodes; at 0.5, it balances.
Improvement: Include a one-shot mode (blinks each time a node is played) and a countdown timer (e.g., time left before 60 seconds for next scene
) in the interface, as in MACAT.
The improved system can:
-
Be trained on a musician’s personal audio corpus in minutes on a CPU, without external GPUs.
-
Generate music that matches the artist’s timbral and stylistic identity, with no risk of mimicking unlicensed artists.
-
Operate in two modes: reactive (follows live input) and proactive (leads with learned patterns), switchable in real time.
-
Self-listen and adapt its output based on audio features (e.g., reduce noisiness, increase tonal clarity).
-
Provide visual feedback (SOM, current/previous node) for explainability and trust.
-
Allow fine-grained control over creativity (congruence) and timbre (resampling, pitch shift, reverse).
-
Support structured performances with scene timers and one-shot visual cues.
These improvements directly enable a human musician to co-create with an AI in live, real-time settings, expanding artistic expression without requiring large-scale data or computational resources.
Abstract
Our research explores the development and application of musical agents, human-in-the-loop generative AI systems designed to support music performance and improvisation within co-creative spaces. We introduce MACAT and MACataRT, two distinct musical agent systems crafted to enhance interactive music-making between human musicians and AI. MACAT is optimized for agent-led performance, employing real-time synthesis and self-listening to shape its output autonomously, while MACataRT provides a flexible environment for collaborative improvisation through audio mosaicing and sequence-based learning. Both systems emphasize training on personalized, small datasets, fostering ethical and transparent AI engagement that respects artistic integrity. This research highlights how interactive, artist-centred generative AI can expand creative possibilities, empowering musicians to explore new forms of artistic expression in real-time, performance-driven and music improvisation contexts.
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning