Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
Junhao Chen, Mingjin Chen, Jingjia Mao, Lin Chen, Saining Zhang, Minglin Chen, Ruocheng Wu, Liaoyuan Fan, Wenyi Li, Mingju Gao, Henghaofan Zhang, Zhihao Li, Hao Zhao, Yufei Wang, Ruqi Huang
cs.SD, cs.CL
Submitted: 2026-08-04
Comments: Project Page: https://yisuanwang.github.io/Agogic
Project page: https://yisuanwang.github.io/Agogic
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Text-to-music language models begin with a choice usually made by default: how to tokenize music.
Terminology
Abstract
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizations, anchoring texture metrics to each representation's model-free ceiling. The ordering is clean and surprising: representation, not model size, is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves Frechet Music Distance (FMD), whereas switching representation halves it. PMT, a performance-resolution stream we release (10 ms timing, per-note velocity, multi-track texture; 609 symbols), reaches FMD 159 at 0.8B against 272-286 for beat grids (1.7-1.8x lower, up to 2.8x elsewhere; non-overlapping bootstrap CIs), so a 0.8B performance-resolution model beats a 27B beat grid. It reappears on a 26M from-scratch backbone and a second performance-resolution tokenizer: a property of the class, not one lucky vocabulary. Nor is it a finer-lattice artifact: snapping PMT's onsets to the beat grids' resolution still leaves it 67-129 FMD ahead of both (n=500). The effect is distributional; whether it is audible is a separate question, left open by our probe, with a human study pre-registered. Native caption adherence is weak but separable: a lightweight decode-time constraint doubles instrument-F1 (.28 to.60) and Correct-Key (.16 to.35) at no distributional cost. We release the harness, 25+ checkpoints, two corpora (86.6k aligned across caption/MIDI/ABC/audio; 6.25M captioned, the largest for music), and an imprinting diagnostic: published text-to-MIDI systems reproduce their training distribution near-invariant to the caption (72% vs. 71% chord-time on disjoint domains). The field's next representation claim can now be measured, not asserted.
Sources
- MusicLM: Generating Music From Text
- Text2Score: Generating Sheet Music From Textual Prompts
- Engine-Native Editable 3D World Reconstruction with Objects and Lighting
- PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation
- Bunraku: Turning a Single Illustration into an Editable Live2D Character
- Soulstyler: Using Large Language Model to Guide Image Style Transfer for Target Object
- One Video, One World: Turning Monocular Video into Physical 4D Scenes
- A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR
- Segment-Factorized Full-Song Generation on Symbolic Piano Music
- Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
- ComposerX: Multi-Agent Symbolic Music Composition with LLMs
- Jukebox: A Generative Model for Music
- MMM : Exploring Conditional Multi-Track Music Generation with the Transformer
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- Emotion-driven Piano Music Generation via Two-stage Disentanglement and Functional Representation
- EmoGen: Eliminating Subjective Bias in Emotional Music Generation
- MuseCoco: Generating Symbolic Music from Text
- Foundation Models for Music: A Survey
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment