UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations

arXiv:2608.00576 · cs.SD, cs.AI, cs.LG · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations".

Jane: The paper was written by Ziyue Kang, Nan Nan, Chenhao Lin and Xiaohong Guan from Xi'an Jiaotong University and Tsinghua University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone! Today we're digging into a paper that's got a mouthful of a title — "UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations." Jane, I have to say, even the title sounds like it needs a translator.

Jane: It really does, Tom. But once you unpack it, it's actually a pretty practical problem. Think about a full orchestral score — you've got dozens of instruments playing at once. Now imagine you need to squeeze that into a format that only allows, say, eight tracks. That's what they're tackling.

Tom: Right, and the authors are from Xi'an Jiaotong University and Tsinghua — Ziyue Kang, Nan Nan, Chenhao Lin, and Xiaohong Guan. They're basically saying that when you compress a big orchestral piece into a smaller format, you can't just chop off the "least important" parts and call it a day.

Jane: Exactly. Because in music, the melody, the bass line, the harmony — they all work together. If you just prune tracks, you might kill the bass line while keeping a filler string part. The paper's argument is that this should be treated as a routing problem — like figuring out which musical material goes where, what gets merged, and what gets dropped.

Tom: And they've built a framework called UOT-IR that does this automatically. It's based on something called unbalanced optimal transport, which is a fancy mathematical way of saying "move stuff around, but you're allowed to lose some of it along the way."

Jane: Which is exactly what you need here. You can't keep everything — the budget is fixed. But you want to keep the right things and put them in compatible slots. It's like packing a suitcase where every item has to fit in a specific compartment, and you have to decide what's worth bringing.

Tom: I love that analogy. And they test this on the SymphonyNet corpus, which is a big collection of orchestral music. They show that their method beats a bunch of simpler baselines — random selection, greedy selection, even some machine learning approaches like PCA and clustering.

Jane: The key insight is that music has structure — instruments have roles, ranges, and timbres that matter. If you ignore that, you get outputs that technically fit the budget but sound wrong or are unplayable. This paper tries to respect that structure while still compressing.

Tom: And that's the part that gets me excited. Because this isn't just about music — it's about any domain where you have rich, multi-part data that needs to fit into a constrained format. But let's not get ahead of ourselves. Next segment, we'll actually walk through what the method does step by step.

Jane: Sounds good, Tom. Stick around, folks — we're just getting started with "UOT-IR."

Summary of the Paper: Jane: So Tom, last segment we set the stage. Now let's actually talk about what UOT-IR does under the hood, because there's some clever stuff here.

Tom: Please, break it down for me. I'm still wrapping my head around "unbalanced optimal transport."

Jane: Okay, so imagine you have a bunch of source tracks — say, violins, trumpets, a piano, a drum kit. And you have a fixed number of target slots — let's say four. The transport part is about moving musical content from the source tracks into those slots. The "unbalanced" part means you're allowed to discard some content — you don't have to preserve everything.

Tom: So it's like a matching problem with permission to say "no" to some matches.

Jane: Exactly. And the clever part is how they decide what matches are good. They use something called Orch2Vec — an orchestration prior that knows which instruments are compatible with each other. It's built from a taxonomy of one hundred twenty-nine instrument tokens and co-occurrence statistics from real music. So a violin track is more likely to be routed to a string slot than to a drum slot.

Tom: That makes sense. But what about the temporal aspect? Music changes over time — a piece might start with just strings and then bring in the full orchestra.

Jane: Right, and that's why they solve this routing problem bar by bar. Each bar gets its own transport plan. But then they need to make sure the slots stay consistent across bars — you don't want the "violin" slot suddenly becoming the "trombone" slot in the next bar for no reason. So they use a temporal decoding step, kind of like a Viterbi algorithm, to keep things stable.

Tom: And they also adapt the "strictness" of the compression per bar. Some bars are dense and need more aggressive pruning, others are sparse and can keep more.

Jane: Precisely. They call it adaptive relaxation — each bar gets its own parameter that controls how much mass is retained. It's not a one-size-fits-all compression.

Tom: And then there's a playability-aware projection at the end — making sure the notes actually fit the target instrument's range and polyphony limits. You can't put a bassoon line into a piccolo slot.

Jane: Right. So the whole pipeline is: build descriptors for each track, compute a cost matrix that combines orchestration compatibility, feature similarity, and physical feasibility, solve the unbalanced transport, decode temporally, and project into playable output.

Tom: It's a lot of moving parts, but the results seem to justify it. In their adaptive preservation setting — where you don't have a fixed template — they hit a Note-F1 of zero point nine one two zero, which is the best among all methods. And in template standardization, they get the lowest structural cost and confusion rate.

Jane: And those numbers matter because they show that the method isn't just keeping notes — it's keeping the right notes in the right places. That's the hard part.

Tom: Alright, I'm convinced this is a real contribution. But I want to know — what does this mean for actual music tools? Let's bring in Meng and Lu for that in the next segment.

Improvements and Implications: Tom: Welcome back, everyone. We've got Meng and Lu with us now. Meng, you're the engineer — what does a paper like this actually enable in practice?

Meng: Well Tom, the first thing that jumps out at me is the training-free aspect. That's huge for deployment. You don't need to train a neural network, you don't need a GPU cluster — you just need the orchestration prior and a solver for the transport problem. That means this could run in real time on a laptop.

Jane: And that's important because a lot of music software — DAWs, notation tools — they all need to handle tracks that exceed their limits. If you're working with a thirty-track orchestral mockup and your notation software only supports sixteen staves, you need a way to reduce it without destroying the music.

Meng: Exactly. And the fact that they handle both template standardization — where you have a fixed target format — and adaptive preservation — where you just keep the best content — means it's flexible enough for different workflows.

Lu: I want to build on that, because I think the implications go beyond just compression. This is really about structured representation learning for music. The routing framework forces you to think about what's essential in a piece — what roles need to be preserved, what instruments are compatible, what's physically playable.

Tom: So you're saying this could inform how we design music generation models?

Lu: Absolutely. Look at NotaGen, which they cite — it excludes scores with more than sixteen staves because of generation complexity. With a method like UOT-IR, you could pre-process those scores into a bounded representation and then feed them into the model. That expands the training data and the creative possibilities.

Jane: And it's not just generation. Think about music analysis — if you can compress a complex score into a fixed number of tracks while preserving structure, you can compare pieces more easily, cluster them, search them.

Meng: I'd add one practical concern though — the orchestration prior is built from co-occurrence statistics in the corpus. That means it's biased toward the styles in that corpus. If you apply this to, say, electronic music or non-Western orchestration, the compatibility scores might be off.

Lu: That's a fair critique. But the framework is modular — you could rebuild the prior from a different corpus. The transport formulation doesn't care where the cost matrix comes from.

Tom: So it's a solid foundation that can be adapted. That's the mark of good research, honestly.

Jane: And there's a cultural angle here too. Lalam, you've been quiet — what do you think about the broader impact?

Lalam: I think the most exciting implication is access. Orchestral music is often locked behind expensive notation software and expert arrangers. A tool like this could let a hobbyist take a full orchestral score and automatically produce a piano reduction or a quartet arrangement that's musically coherent. That's a democratization of arrangement.

Tom: That's a beautiful way to put it. So we've got practical engineering value, research implications for generation and analysis, and cultural access. Not bad for a paper about "routing."

Jane: Let's wrap up with a final summary in the next segment.

Conclusion: Tom: Alright, we're at the finish line. Let's pull it all together for "UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations."

Jane: So the core idea is simple: when you need to compress a multi-track piece into a fixed number of slots, treat it as a structured routing problem. Don't just prune tracks or project into a lower-dimensional space — explicitly decide which musical material goes where, what gets merged, and what gets dropped.

Tom: And they do that with unbalanced optimal transport, which allows selective discard, combined with an orchestration prior, temporal coherence, and playability checks. It's a complete pipeline.

Jane: The results on SymphonyNet show it beats heuristic baselines and generic representation-space methods on both fidelity and structural compatibility. The full model gets the best Note-F1 in adaptive preservation and the lowest structural cost in template standardization.

Meng: And from an engineering standpoint, it's training-free, which means it's practical to deploy. You could integrate this into existing music software without a huge infrastructure investment.

Lu: And from a research standpoint, it opens up new ways to think about bounded symbolic representations — not as a limitation, but as a design choice that can be optimized.

Lalam: And culturally, it makes orchestral arrangement more accessible to people who aren't professional orchestrators.

Tom: That's a lot of value from one paper. I think we can safely say this is one to watch.

Jane: Agreed. We'll be keeping an eye on follow-up work — especially if they extend this to other musical styles or integrate it into generation models.

Tom: Thanks for joining us, everyone. That's it for "UOT-IR." Next up, we've got a paper on neural audio codecs — so stay tuned.

Jane: See you then, folks!

Ziyue Kang, Nan Nan, Chenhao Lin, Xiaohong Guan

Xi'an Jiaotong University · Tsinghua University

cs.SD, cs.AI, cs.LG

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: 8 pages, 2 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 66/100

Key concepts

UOT-IR
This is a framework that structures the process of compressing high-polyphony symbolic music into a fixed budget. It uses unbalanced optimal transport to decide which musical material to keep, merge, or drop based on compatibility and structure.
Unbalanced Optimal Transport
This mathematical concept allows for moving content from source tracks into target slots while permitting the discarding of some content. It is used here because the budget (fixed number of slots) cannot be exceeded, necessitating selective pruning.
Orchestration Prior
This prior is a mechanism that knows which instruments are compatible with each other. It is built from a taxonomy of instrument tokens and co-occurrence statistics from real music to guide the routing process toward musically sensible results.

Terminology

Summary

Summary

This paper introduces UOT-IR (Unbalanced Optimal Transport for Information Routing), a training-free framework for compressing high-polyphony symbolic music into fixed-budget representations. The authors state: "To address the issue, this study reformulates the compression problem as a fixed-budget structured routing problem and proposes Unbalanced Optimal Transport for Information Routing (UOT-IR), a training-free framework based on constrained unbalanced optimal transport."

The problem is motivated by the fact that "many symbolic music representations and processing frameworks impose fixed budgets on the number of tracks, parts, or target slots [1–4]. Once a richly orchestrated score exceeds such a structural budget, it must first be converted into a bounded form. Existing approaches fall into two categories: heuristic simplification (e.g., pruning, merging, rule-based selection) and representation-space reduction (e.g., PCA, NMF, K-means). The authors argue these often fail to preserve structural roles, orchestration compatibility, and playability under strict budgets, leading to three practical failure modes: salient lines may be removed, incompatible materials may be merged into the same bounded slot, and the resulting output may violate basic playability or role-consistency requirements."

The method is described as follows: "UOT-IR is a training-free framework based on constrained unbalanced optimal transport that integrates a taxonomy-grounded orchestration prior, symbolic statistical descriptors, adaptive marginal relaxation, temporal coherence, and playability-aware projection to produce compact yet musically coherent bounded outputs." The framework operates on bar-level segmentation, where for each bar b, it solves for a nonnegative transport matrix(b) in R N b times K+, with N b active source tracks and K target slots. The optimization problem is:

[

(b) at least 0(b), C(b) + lambda s D rho b((b) 1, mu(b)) + lambda t D rho b(((b)) 1, nu(b))

]

where C(b) is the routing cost matrix, mu(b) and nu(b) are source and target marginals, and D rho b is the generalized Kullback–Leibler divergence with relaxation parameter rho b.

The routing cost matrix is decomposed into four components: C(b) = alpha C prior + beta C local + gamma C semantic + delta C physical. The prior term uses Orch2Vec, a program-level orchestration prior built from a 129-instrument vocabulary (128 General MIDI programs plus a drums token) organized in a tree-structured taxonomy. This is combined with data-driven co-occurrence statistics using positive pointwise mutual information (PPMI) to form a symmetric prior matrix C prior in R 129 times 129. The local term measures cosine distance between bar-level source descriptors and target prototypes. The semantic term captures role-related cues from track and slot names. The physical term penalizes pitch-range incompatibility.

The framework includes three key refinements: (1) adaptive relaxation, where UOT-IR therefore uses a bar-specific relaxation parameter rho b chosen by matching retained mass to a desired target level via rho b = rho in R (b)(rho) 1 - tau b; (2) temporal decoding via a target-aware sticky decoding scheme, implemented as a Viterbi-style dynamic program with transition cost T(a,b) = lambda stay 1[a not equal to b] + lambda prog 1[prog(a) not equal to prog(b)]; and (3) playability-aware projection including pitch-range correction, short-note filtering, and target-dependent polyphony control.

The paper studies two settings under the same slot budget: template standardization, which maps each input to a predefined bounded target template, and adaptive preservation, which preserves representative content without assuming an external template. Experiments are conducted on the SymphonyNet corpus, with evaluation subsets constructed from pieces whose active track count exceeds the target budget.

Baselines include heuristic methods (Direct, Random, Greedy, Skyline, BMF-PC), representation-space methods (PCA, NMF, KMeans), and transport baselines (Vanilla-UOT, UOT-IR-Core). Metrics cover content fidelity (Note-P, Note-R, Note-F1, FTED, PC-JSD, Dur-JSD, IOI-JSD) and structural compatibility (SC, BC, PR-Diff, PCE-Diff).

Key results: In adaptive preservation (Table 1), UOT-IR achieves the best overall Note-F1 and recall with Note-F1 of 0.9120 and Note-R of 0.9200. In template standardization (Table 2), UOT-IR achieves the best overall task-relevant performance, especially on PR-Diff, PCE-Diff, SC, and BC, with SC of 14.7165 and BC of 0.3406. The ablation study (Table 3) shows that the prior improves structural compatibility, while UOT and TTA are important for adaptive preservation; overall, the full model achieves the best balance across metrics.

The authors conclude: "This paper shows that over-budget symbolic music is better modeled as a structured routing problem than as heuristic simplification or generic representation-space reduction. UOT-IR is a training-free framework based on constrained unbalanced optimal transport. It unifies template standardization and adaptive preservation within a single formulation. They note that Future work will study bounded routing-based representations for controllable arrangement, orchestration-aware generation, and fixed-budget symbolic music modeling."

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Implement UOT-IR as a preprocessing layer for any music generation or analysis model that requires bounded track/slot representations.

What the improved system can do:

  • Accept high-polyphony symbolic scores (e.g., 20+ tracks) and automatically compress them to any fixed budget (e.g., 8 slots) while preserving musical structure

  • Achieve Note-F1 of 0.9120 in adaptive preservation and 0.9370 in template standardization, outperforming heuristic baselines by 3–5%

  • Reduce structural cost by 22% compared to best heuristic baseline (SC: 14.72 vs. 18.88 for BMF-PC)

  • Handle both template standardization (mapping to predefined orchestration templates) and adaptive preservation (retaining representative content without templates)

Abstract

High-polyphony symbolic music is increasingly used in generation, analysis, and arrangement, yet many downstream tasks require bounded representations with fixed tracks or slots. Converting richly orchestrated scores into compact forms is therefore necessary, but existing approaches relying on heuristic simplification or generic representation-space reduction often fail to preserve structural roles, orchestration compatibility, and playability under strict budgets. To address the issue, this study reformulates the compression problem as a fixed-budget structured routing problem and proposes Unbalanced Optimal Transport for Information Routing (UOT-IR), a training-free framework based on constrained unbalanced optimal transport. UOT-IR combines an orchestration prior, adaptive marginal relaxation, temporal decoding, and playability-aware projection to produce compact and musically coherent bounded representations. This work further studies two practical settings under the same slot budget: template standardization, which maps each input to a predefined bounded template, and adaptive preservation, which retains representative content without assuming an external template. Experiments on the SymphonyNet corpus show that UOT-IR delivers strong overall performance across both settings, including the best Note-F1 in adaptive preservation (0.9120), together with the lowest structural cost (14.7165) and bad structural confusion rate (0.3406) in template standardization. This work establishes a principled paradigm for fixed-budget symbolic music compression, offering a practical path toward compact, structured, and musically coherent symbolic representations.

Sources

Related papers