Learning and composing of classical music using restricted Boltzmann machines
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Learning and composing of classical music using restricted Boltzmann machines".
Jane: The paper was written by Mutsumi Kobayashi and Hiroshi Watanabe from Keio University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making the rounds on arXiv, and it’s called “Learning and composing of classical music using restricted Boltzmann machines.” Jane, I have to say, the title alone brings me back to my old machine learning classes.
Jane: It does, Tom, and that’s exactly why I love it. So, a restricted Boltzmann machine, or RBM for short, is basically a simple type of neural network that learns patterns from data. The authors here, Mutsumi Kobayashi and Hiroshi Watanabe from Keio University, they took piano roll images of Bach’s music and fed them into this RBM.
Tom: Right, and a piano roll is that old-school player piano format, right? The paper converts musical scores into these black-and-white images where time runs horizontally and pitch runs vertically. So the model is essentially learning to look at music like it’s a picture.
Jane: Exactly. And the big deal here is that they wanted to see if this really simple model could compose music, and more importantly, what it actually learns inside. They’re not trying to build the next big AI composer, they want to open up the black box.
Tom: And that’s what gets me excited. They’re asking, “Can a minimal model generate something musical, and can we peek inside to see how it’s doing it?” That’s a really refreshing approach compared to the giant deep learning models we usually hear about.
Jane: Totally. And the paper’s authors are pretty upfront that they’re prioritizing interpretability over sheer generative performance. They want to use the model as a tool to understand musical structure, not just to churn out pieces.
Tom: So, for our listeners, this is a paper about using a very transparent, simple AI to compose classical music, and then dissecting that AI to see what musical concepts it picked up on its own. It’s like giving a kid a box of LEGOs and then asking them to explain their building rules afterwards.
Jane: That’s a great way to put it. And the implications are pretty big for how we think about AI in creative fields. If a simple model can learn musical structure, what does that say about what’s necessary for creativity? We’ll get into the actual results in a moment.
Tom: Stick around, because we’re about to see if this little RBM can actually hold a tune.
Paper discussion segment 2: Jane: Alright, we’re back, and we’re still on “Learning and composing of classical music using restricted Boltzmann machines.” So, Tom, we set the stage, but what did the authors actually find?
Tom: Well, Jane, the first thing they did was test if the model could reconstruct piano rolls. They fed it a Bach piece it had seen before, and it rebuilt it perfectly. Then they fed it a Mozart piece it had never seen, and it still reconstructed it accurately. That’s a good sign that it learned general musical features, not just memorized the training data.
Jane: And here’s the kicker—they fed it MNIST digit images, you know, the handwritten numbers. The model completely failed to reconstruct those. It just produced noise. So the RBM definitely knows what a piano roll looks like, and it knows what it doesn’t look like.
Tom: That’s such a clean experiment. They even computed the energy of the model for different inputs. Piano rolls had really low energy, like negative three thousand six hundred while the digit images and noise had much higher, sometimes positive, energy. It’s like the model has a built-in “musicality” meter.
Jane: Right, and lower energy in a Boltzmann machine means the model considers that input to be more probable. So it’s essentially saying, “Yes, this looks like the music I learned, and that digit image looks like nonsense to me.”
Tom: Then comes the fun part—they actually had it compose. They developed an algorithm to generate new piano rolls. The two-measure pieces it created were surprisingly coherent. The paper notes that the generated segment was in B minor, contained a proper E minor chord, and even had a nice stepwise melody.
Jane: And that’s the part that made me sit up. They weren’t just generating random notes; the model was producing something that follows musical rules, like harmony and key. But then they tried to extend it to eight measures, and that’s where things fell apart.
Tom: Yeah, the longer pieces started out okay, in F major, but after a few measures the pitch content got messy and the musical coherence just dissolved. So the model can handle two measures, which is what it was trained on, but it can’t really hold a longer narrative together.
Jane: That’s a really honest result. It shows the limits of the model, but it also tells us something important. The RBM learned local musical grammar, like chords and short melodic phrases, but it didn’t learn long-term structure. It’s like it knows how to write a good sentence but not a good paragraph.
Tom: And that’s a huge insight for anyone building music AI. It suggests that long-range structure needs a different mechanism, maybe something with memory. But for a simple model, getting the local grammar right is already pretty impressive.
Jane: Exactly. And this sets us up perfectly for the next part, where we talk about what the authors found when they looked inside the model’s brain, so to speak.
Paper discussion segment 3: Tom: Welcome back. We’re still on “Learning and composing of classical music using restricted Boltzmann machines,” and now we get to the part I’ve been waiting for—the internal representations. Jane, what did they see when they poked around inside?
Jane: So, they did something really clever. They activated individual hidden units in the model, one at a time, and looked at what kind of visible pattern that unit was responsible for. It’s like asking each neuron, “What are you looking for in the music?”
Tom: And what did they find? I’m guessing it wasn’t a perfect “this is a C major chord” detector.
Jane: You’d be right. What they found were a lot of local temporal patterns, mostly related to note duration. The hidden units had learned to recognize things like sixteenth-note rhythms. So the model was really good at understanding rhythm, which makes sense because that’s a very visual, structural feature of a piano roll.
Tom: But they didn’t find anything that looked like a melodic phrase or a chord shape, right? The paper says those were barely observed. So the model’s internal language isn’t the same as our music theory language.
Jane: Exactly. And that’s a really profound finding. The model is learning the statistical structure of the data, but it’s not learning it in a way that maps neatly onto human concepts like “dominant seventh chord” or “cadence.” It’s a data-driven representation that’s fundamentally different from how we think about music.
Tom: And that has huge implications for explainable AI. We want to trust these models, but if their internal logic is alien to us, how do we build that trust? The authors are essentially saying, “Hey, we looked, and it’s not interpretable in the way we hoped.”
Jane: Right. But they also suggest this isn’t necessarily a bad thing. The latent space of the RBM could be used as a new analytical tool. It’s a way to look at music that isn’t bound by traditional theory. Maybe it can find patterns we haven’t noticed.
Tom: That’s a really creative spin on it. Instead of forcing the model to explain itself in our terms, we learn to read its terms. And the paper also points out that this lack of interpretability might be because the RBM’s representations are complex mixtures of features. They mention that classification RBMs might produce more distinguishable units.
Jane: So there’s a clear path forward. They’re not just saying, “It’s a black box, give up.” They’re saying, “Here’s what we found, and here’s how we might build a model that’s more interpretable.” That’s the kind of constructive criticism we need in this field.
Tom: And it makes me wonder about the bigger picture. If a simple model learns rhythm but not harmony in a human-readable way, what does that say about how we teach music? Maybe our theoretical framework is just one way to slice the pie.
Jane: That’s a great question, Tom, and I think we should bring in the rest of the team to chew on that. But first, let’s wrap up our thoughts on this paper.
Conclusion: Jane: Alright, we’re wrapping up our discussion on “Learning and composing of classical music using restricted Boltzmann machines.” Tom, give us the final takeaway.
Tom: Sure, Jane. The paper showed that a simple RBM can learn to compose two-measure pieces of classical music that are musically coherent, at least locally. It can even distinguish between musical and non-musical images. But when they tried to generate longer pieces, the coherence fell apart.
Jane: And the most important finding was that the model’s internal representations don’t match human music theory. It learned rhythm patterns, but not recognizable chords or melodies. That tells us that a minimal generative model can capture statistical regularities, but its way of organizing that information is fundamentally different from ours.
Tom: Right. And that’s a big deal for the field of explainable AI in creative tasks. It’s a reminder that just because a model can produce something that sounds good, doesn’t mean we can easily understand how it’s doing it. The authors did a great job of being honest about that limitation.
Jane: And they also gave us a roadmap for the future. They suggested looking at different composers, analyzing the weight matrices, and trying different architectures like classification RBMs. So this isn’t a dead end, it’s a starting point for building more interpretable music AI.
Tom: Absolutely. And on a personal note, I love that they made their code available on GitHub. That’s how science moves forward—by letting others build on your work.
Jane: Couldn’t agree more. So, to the authors, Kobayashi and Watanabe, thank you for this thoughtful paper. To our listeners, we hope you enjoyed this deep dive. We’re going to say goodbye to this paper and get ready to look at the next one.
Tom: See you all next time, and keep listening to the music of the data.
Jane: Take care, everyone.
Mutsumi Kobayashi, Hiroshi Watanabe
Keio University
cs.SD, cs.LG, eess.AS
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 19 pages, 12 figures, manuscript was revised
Code: https://github.com/watanabe-appi/simple_rbm
Project page: https://watanabe-appi.github.io/rbm-music-demo
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 67/100
The gist: This study investigates how machine learning models acquire the ability to compose music and how musical information is internally represented within such models.
Key concepts
- Restricted Boltzmann Machine (RBM)
- A simple type of neural network that learns patterns from data. In this study, it was used to analyze piano roll images of classical music to see if a minimal model could learn musical structure.
- Piano Roll Images
- Black-and-white images created by converting musical scores. Time runs horizontally and pitch runs vertically, allowing the model to look at music like a picture.
- Interpretability
- The authors prioritized understanding what the model learns inside its structure over just achieving high generative performance. The hosts discuss how this lack of interpretability relates to building trust in AI.
Terminology
Summary
This study investigates how machine learning models acquire the ability to compose music and how musical information is internally represented within such models. The authors develop a composition algorithm based on a restricted Boltzmann machine (RBM), a simple generative model capable of producing musical pieces of arbitrary length. They convert musical scores into piano-roll image representations and train the RBM in an unsupervised manner. They confirm that the trained RBM can generate new musical pieces; however, by analyzing the model’s responses and internal structure, they find that the learned information is not stored in a form directly interpretable by humans. This study contributes to a better understanding of how machine learning models capable of music composition may internally represent musical structure and highlights issues related to the interpretability of generative models in creative tasks.
The authors deliberately focus on generative models with transparent and straightforward structures, whose internal states are more amenable to systematic analysis. Rather than aiming to develop a highly optimized composition system, they seek to design an algorithm that enables music generation with a minimal, interpretable model architecture. The central objective is to investigate how such a simple model can acquire the ability to generate music, and to analyze its responses and internal representations once this ability has emerged. By prioritizing interpretability over sheer generative performance, this study adopts a constructive perspective aligned with the goals of explainable AI, treating the generative model not only as a creative system but also as a tool for probing the internal mechanisms of machine learning-based musical representation.
The authors adopt the Restricted Boltzmann Machine (RBM) as their modeling framework, as its simple, constrained architecture is well suited to their goal of analyzing internal representations transparently. An RBM is a probabilistic generative model composed of a visible layer and a hidden layer, and it learns the probability distribution of training data. In contrast to general Boltzmann machines, the RBM imposes restrictions on network connectivity, resulting in a simpler learning algorithm and a set of internal parameters that are more directly amenable to systematic analysis. Although a standard RBM cannot explicitly model the temporal structure of input sequences, and extensions such as the Temporal RBM (TRBM) and Conditional RBM (CRBM) have been proposed to address this limitation, the authors deliberately refrain from introducing such temporal mechanisms. By doing so, they prioritize interpretability and generate music strictly within the standard RBM framework.
The authors train a restricted Boltzmann machine (RBM) on piano-roll images derived from keyboard works by J. S. Bach and conduct a multifaceted analysis of its behavior. Specifically, (i) they evaluate how accurately the trained RBM reconstructs the training piano-rolls and analyze the energy values it assigns to unseen piano-rolls and to non-musical images, such as MNIST digits, thereby assessing its ability to distinguish musical from non-musical data. Furthermore, (ii) they generate new piano-roll samples from the trained RBM using the proposed generation algorithm and examine the extent to which it can produce musically coherent structures in sequences of two measures or longer. In addition, focusing on the internal representations of the RBM, (iii) they input one-hot vectors to individual hidden units and analyze the expected visible-layer patterns to investigate the types of musical patterns encoded by each hidden unit.
Through these analyses, the authors aim to clarify what kinds of musical regularities a standard RBM, here used as a minimal generative model, can learn from piano-roll images of J. S. Bach’s compositions, and whether its internal representations correspond to concepts familiar to human music theory. The results demonstrate that RBMs are capable of musical generation and further suggest that the latent space learned by an RBM may function as a data-driven analytical representation that is not necessarily aligned with conventional music-theoretical frameworks.
For the training data, the authors adopted compositions by J. S. Bach. A total of 58 MIDI files were obtained from the Mutopia Project, and each file was converted into a black-and-white image representation known as a piano roll. A piano roll is a two-dimensional representation of musical information, where the horizontal axis corresponds to time and the vertical axis corresponds to pitch. Notes are depicted as horizontal bars, with their positions and lengths indicating the timing and duration of each note, respectively. Each pixel value in the piano roll image is either 0 or 1, corresponding to the binary visible units of a Bernoulli-type RBM. In order to standardize the input dimensions, the training data was restricted to compositions in 4/4 time. The musical sequences were then partitioned so that each image corresponded to two measures of music. The image size was fixed at 72×192 pixels. The vertical dimension of 72 pixels corresponds to the pitch range from C1 to B6, where C1 denotes the C note in the first octave of a standard 88-key piano (i.e., the lowest C key), and B6 denotes the B note in the sixth octave, one semitone below the highest C (C8). The horizontal dimension of 192 pixels represents time, with 24 pixels corresponding to the duration of one quarter note. This resolution was chosen so that the horizontal pixel count would be divisible by 3, enabling the representation of triplet notes. During training, each piece was transposed into a total of 11 keys, including keys up to 6 semitones higher and 5 semitones lower than the original key. Through this process, a dataset of 22,116 images for training was obtained. The image size was 72 × 192 and the number of hidden units was 2048.
The authors composed music using an RBM trained on piano rolls of compositions by J. S. Bach. The number of visible units in the RBM used for training was 13,824, which corresponds to two measures in 4/4 time. Therefore, the above method allows the model to compose music up to a maximum length of two measures. To enable the RBM to generate music longer than two measures, the authors adopted an iterative procedure in which the latter one measure of the generated two-measure sequence are fixed and used as the first one measure for the next generation step. By repeating this procedure, the RBM is able to generate longer musical sequences. The composition procedure is as follows: (1) Initialize all visible units to zero and denote the resulting vector as v0. (2) Set vt as the visible state of the trained RBM. (3) Given the visible units fixed at vt, the hidden unit states are sampled using Gibbs sampling. (4) Compute the expected visible state ut given the sampled hidden unit states fixed. (5) Construct the next visible vector vt+1 by setting the t + 1 largest elements of ut to 1 and the rest to 0. Note that elements which were 1 in vt may become 0 in vt+1. (6) By repeating steps 2 through 5 N times, a binary vector is obtained in which exactly N elements are set to 1. The authors set N = 1000 for generating the initial two measures, and N = 500 for the process in which the right half of a measure is generated while keeping the left half fixed. This extension process was repeated six times, and the resulting images were concatenated to produce a final piano roll corresponding to eight measures of music.
To verify whether the trained RBM correctly memorized the piano rolls, the authors input the piano roll into the visible units and examined whether it could be reconstructed through Gibbs sampling. First, when a piano roll of a J. S. Bach composition used during training was provided as input, the RBM successfully reconstructed it. The authors also provided a piano roll of a W. A. Mozart composition that was not included in the training data. The RBM was still able to reconstruct the image. From these results, the authors conclude that the RBM has acquired the capability to accurately reconstruct piano roll images. To evaluate whether the RBM trained on piano rolls can reconstruct images outside the training domain, the authors used the MNIST dataset as input. Each 28 × 28 pixel image was resized to 72 × 192 pixels and provided to the visible units. In contrast to the case of piano roll images, the RBM failed to reconstruct digit images and instead produced noise-like outputs. These results indicate that the RBM trained on piano rolls is capable of reconstructing unseen piano roll images, but not images that differ in nature, such as handwritten digits. This confirms that the RBM has learned the specific features of piano roll images.
To investigate how the energy of the trained RBM responds to piano roll images versus non-piano roll images, the authors input various types of images into the RBM and computed the corresponding energy values. Specifically, for each image, the corresponding binary vector was fed into the visible units, and the hidden units were sampled using Gibbs sampling. The energy of the RBM was then calculated from the visible and hidden states. As input images, the authors used a piano roll included in the training data, a piano roll not used during training, three digit images from the MNIST dataset (0, 5, and 8), and white noise. They determined averages and standard deviations from 10 independent samples. As a result, piano roll images exhibited low energy values regardless of whether they were included in the training data, while other types of images generally resulted in positive energy values. Although some MNIST digit samples showed negative energy, their values were still significantly higher than those of the piano roll images. These results indicate that the RBM has learned to assign lower energy to visible unit configurations resembling piano rolls. The energy values were: Piano roll (trained) −3654 ± 4; Piano roll (untrained) −3353 ± 3; MNIST digit 0 44.8 ± 0.4; MNIST digit 5 −443.4 ± 1.8; MNIST digit 8 −0.1 ± 0.7; Noise 83.7 ± 0.1.
An example of two-measure music generation using the composition algorithm is shown. The figure shows the visible states vt at sampling steps t = 50, 100, 250, 500, 750, and 1000. All images exhibit the structure of piano rolls. The time evolutions of the energy of the RBM during image generation shows that the energy decreases monotonically up to approximately 500 sampling steps, after which it begins to increase. This suggests that the RBM assigns higher energy when the number of active pixels (notes) is either too small or too large, implying the existence of an optimal number of notes that minimizes the energy. The energy reached its minimum at sampling step t = 557. An analysis of this piano roll reveals that all notes appearing in the segment belong to the pitch-class set of B minor. In addition, the diatonic chord E minor, which is one of the diatonic triads in B minor, is present in the generated segment. The phrase also contains a stepwise motion C#-D-C#-B, which is musically natural in the context of the B-minor scale. These observations indicate that the piano roll generated by the RBM exhibits musically ordered structure in terms of overall pitch content, harmonic organization, and melodic motion. An example of an eight-measure composition generated using the extended algorithm also exhibited a piano roll structure, similar to the two-measure images. A close inspection of the piano roll shows that the pitch organization in measures 1-3 is based on the F-major key, exhibiting a musically ordered structure. However, after the third measure, the pitch content gradually becomes more irregular, and the musical coherence diminishes. Therefore, it is considered difficult for the trained RBM in its current form to generate piano rolls that exceed the number of measures in the training data while maintaining musically coherent structure.
To investigate what kinds of patterns the RBM extracted from the musical training data, the authors provided one-hot vectors to the hidden layer of the trained model and computed the corresponding expected values of the visible layer. When these expected values were visualized as a colormap, numerous local temporal patterns with the width of a sixteenth-note duration were observed. This result indicates that the RBM spontaneously extracts elements corresponding to note duration from the input music. Therefore, the RBM can be regarded as having acquired internal representations that enable the reconstruction of rhythm structures based on note values. On the other hand, typical melodic phrases or chordal structures were scarcely observed in the extracted patterns, suggesting that the internal representations of the trained RBM are not readily interpretable in terms of human musical intuition. It has been pointed out that the latent representations of standard RBMs often consist of complex mixtures of multiple features, and it is therefore difficult for the model to acquire feature-separated internal representations—such as those corresponding to specific chordal or harmonic structures—without explicit label information. Nevertheless, the fact that the RBM’s internal representations do not directly correspond to human music-theoretical concepts suggests that the model captures the statistical structure of musical data from a perspective fundamentally different from that of human music theory. In this sense, the latent space acquired by the RBM may serve as a data-driven analytical representation that does not rely on conventional theoretical frameworks.
The authors demonstrated that music composition is feasible even with a structurally simple model such as an RBM. By representing musical scores in piano-roll format, they enabled the model to learn musical features using techniques analogous to those employed in image modeling. The trained RBM successfully reconstructed piano-roll representations, including those derived from musical pieces not seen during training, while failing to reconstruct non-musical images and assigning high energy values to such inputs. Although the training data were limited to two-measure piano rolls, the authors developed a generation algorithm that allowed the model to produce musical sequences of arbitrary length. The simplicity of the RBM architecture allowed the authors to analyze how the trained model internally represents musical data in a more direct manner than would be feasible with more complex models. By examining the hidden-layer activations in response to various inputs, they found that musical transposition caused substantial changes in the internal states, suggesting that the RBM evaluates musical similarity primarily based on the overlap of absolute pitch positions rather than on abstract melodic structure. This behavior is consistent with previous observations that RBMs and Deep Belief Networks lack inherent translational invariance in their input space. In contrast, convolutional deep belief models, which incorporate local receptive fields and weight sharing, offer a potential path toward improved recognition of transposed musical patterns due to their translational invariance.
Future work will extend the present framework to musical corpora beyond the works of J. S. Bach in order to examine whether training on different composers or musical genres leads to systematically distinct generative characteristics. Such studies may clarify whether restricted Boltzmann machines can extract and reproduce composer-specific or genre-specific stylistic features. In addition, recent studies have suggested that the tasks learned by RBMs are reflected in the singular value spectrum of their weight matrices, and examining how the composer or genre influences this spectrum represents a promising direction for further research. Moreover, previous work has indicated that hidden units in RBMs can encode prototypical patterns in the visible layer, although such patterns are often difficult to interpret in standard RBMs. Architectures such as classification RBMs, which tend to produce more distinguishable hidden-unit activations, may facilitate the identification of prototypical melodic or harmonic structures, and exploring such architectural extensions remains an important topic for future investigation.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement and what the improved AI system can do:
- Add pitch-transposition invariance to the RBM architecture
-
Implement a convolutional RBM (as suggested by Lee et al., 2009) with local receptive fields and weight sharing across the pitch axis.
-
This directly addresses the paper’s finding that the standard RBM lacks translational invariance, causing transposed musical inputs to be treated as unrelated.
-
The improved system will recognize melodic and harmonic patterns regardless of key, enabling more robust composition and analysis.
- Introduce a temporal conditioning mechanism without sacrificing interpretability
-
Replace the naive iterative extension (Algorithm 2) with a conditional RBM that explicitly conditions on the previous measure’s hidden state, while keeping the visible–hidden connectivity simple.
-
This maintains the model’s transparency while allowing coherent generation beyond two measures, directly fixing the observed degradation in musical coherence after measure 3.
- Add a post-hoc interpretability layer using sparse hidden-unit activation
-
Apply a sparsity penalty (e.g., L1 regularization on hidden activations) during training to encourage each hidden unit to encode a single, separable musical feature (e.g., a specific chord, rhythmic pattern, or melodic contour).
-
This addresses the paper’s finding that hidden units encode complex mixtures of features, making internal representations more aligned with human music-theoretic concepts.
- Implement a key-aware energy function
-
Modify the energy function to include a global pitch-class bias vector that is learned per transposition level, effectively normalizing for absolute pitch.
-
This will lower energy for musically equivalent transpositions, improving both reconstruction of unseen transposed pieces and generation quality.
- Add a coherence-aware sampling schedule
-
During generation, monitor the energy landscape and the number of active notes per measure; use a feedback loop to adjust the sampling temperature or the number of Gibbs steps per segment.
-
This prevents the energy from increasing after 500 steps (as observed in Fig. 8) and maintains musical density within a natural range.
-
Generate musically coherent compositions of 8+ measures that maintain key stability and harmonic progression, not just the first 2–3 measures.
-
Recognize and reproduce melodic/harmonic patterns across all 12 keys, so a melody learned in C major can be generated in F# major without retraining.
-
Provide human-interpretable internal representations: each hidden unit corresponds to a recognizable musical element (e.g., a B-minor chord, a sixteenth-note rhythmic motif, a stepwise melodic contour), enabling musicians to inspect and steer the generation process.
-
Reconstruct and analyze unseen piano rolls from any composer with higher fidelity, because the model no longer overfits to absolute pitch positions.
-
Serve as an analytical tool for music theory: by examining which hidden units activate for different musical segments, researchers can discover data-driven harmonic and rhythmic rules that may complement or challenge traditional theory.
-
Maintain computational efficiency while being more interpretable, because the convolutional and conditional extensions preserve the RBM’s simple learning rules (CD) and do not require deep architectures.
-
Support controlled composition: by clamping specific hidden units (e.g., those encoding a dominant chord or a syncopated rhythm), the system can generate music with user-specified stylistic constraints.
Sources
- MusicLM: Generating Music From Text
- Exploring XAI for the Arts: Explaining Latent Space in Generative Music
- Jukebox: A Generative Model for Music
- Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset
- C-RNN-GAN: Continuous recurrent neural networks with adversarial training
- Learning Interpretable Representation for Controllable Polyphonic Music Generation
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment