CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling
cs.SD, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-18
Code: https://github.com/FEAfeatherTHER/CPR_official
License: http://creativecommons.org/licenses/by/4.0/
The gist: Prompt-conditioned piano MIDI-to-Music rendering aims to faithfully render target notes while reproducing the timbre of a reference recording.
Terminology
Abstract
Prompt-conditioned piano MIDI-to-Music rendering aims to faithfully render target notes while reproducing the timbre of a reference recording. Existing approaches primarily follow two paradigms: autoregressive (AR) modeling and flow matching (or diffusion). Discrete-codec AR models provide causal temporal modeling, but quantization can discard acoustic detail. Flow matching better preserves acoustic structure in the cost of full-sequence attention costs and worse semantic structure. Continuous autoregressive models operate directly on continuous representations. It not only combines the condition-following ability of AR models and distribution-modeling capacity of flow matching but also bypasses the quantization bottleneck with lower computational costs. Building on this principle, we present Composer--Performer--Refiner (CPR) framework. Composer autoregressively predicts continuous hidden states, Performer generates 24kHz acoustic latents through local flow matching and Refiner then upsamples the waveform to 48 kHz. We further introduce Bottlenecked Representation Alignment (BREPA) and Modality--Time RoPE (MT-RoPE) to strengthen musical semantic structure in Composer hidden states and temporal alignments across modalities. Codes are available at https://github.com/FEAfeatherTHER/CPR official
Sources
- P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- Making Reconstruction FID Predictive of Diffusion Generation FID
- Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
- Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
- Qwen3 Technical Report
- Fr'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment