Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model
cs.SD, cs.LG
Submitted: 2026-05-20
Updated: 2026-09-11
Comments: 32 pages, 13 figures
Code: https://github.com/craffel/mididataset
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: This study aims to enhance the quality of music generation using Transformers by incorporating meta-information.
Terminology
Abstract
This study aims to enhance the quality of music generation using Transformers by incorporating meta-information. While Transformer-based approaches are effective at capturing long-term dependencies in musical compositions, the music they generate often suffers from issues such as excessive repetition or duplication of notes, leading to unnatural melodies. To address these limitations, we propose Musical Attention, a mechanism that incorporates meta-information such as bar numbers, key, signatures, and tempos into the attention process. Musical Attention explicitly leverages both the structural properties of music and its associated metadata, enabling the Transformer's attention mechanism to operate more effectively and thereby improving the quality of the generated output. In our framework, each musical note is represented as a combination of five events-pitch, bar number, onset, duration, and velocity in addition to the three metadata elements. The attention mechanism is then modified to reflect the correlations among these eight features, allowing the model to better capture the inherent characteristics of musical composition. Experimental results demonstrate that the model incorporating Musical Attention outperforms prior methods, such as Full Attention and Strided Attention, in terms of musical coherence, variation, and overall quality. Notably, it significantly reduces repetition and enhances the model's ability to generate diverse, harmonically consistent melodies. Musical Attention thus represents a meaningful advancement in AI-driven music generation, facilitating the creation of more natural and expressive compositions.
Sources
- A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music
- MIDI-VAE: Modeling Dynamics and Instrumentation of Music with Applications to Style Transfer
- Symbolic Music Genre Transfer with CycleGAN
- Attention Is All You Need
- MusicBERT: Symbolic Music Understanding with Large-Scale Pre-Training
- BERT-like Pre-training for Symbolic Piano Music Classification Tasks
- Simple and Controllable Music Generation
- Self-Attention with Relative Position Representations
- Music Transformer
- MusicLM: Generating Music From Text
- MuLan: A Joint Embedding of Music Audio and Natural Language
- MidiNet: A Convolutional Generative Adversarial Network for Symbolic-domain Music Generation
- MuseGAN: Multi-track Sequential Generative Adversarial Networks for Symbolic Music Generation and Accompaniment
- Generative Adversarial Networks
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano Compositions
- MuseMorphose: Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE
- Generating Long Sequences with Sparse Transformers
- Language Models are Few-Shot Learners
- Training language models to follow instructions with human feedback
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment