Architecture and Affordances of PLAUD: Performative Latents and Unsupervised DDSP
cs.SD, cs.HC, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-13
Comments: AI Music Creativity 2026
Code: https://github.com/acids-ircam/nn_tilde
Project page: https://blazejkotowski.com
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: PLAUD (Performative Latents and Unsupervised DDSP) is a neural synthesizer and Max for Live instrument for live electronic music, built on NoiseBandNet and trained on small personal sound corpora.
Terminology
Abstract
PLAUD (Performative Latents and Unsupervised DDSP) is a neural synthesizer and Max for Live instrument for live electronic music, built on NoiseBandNet and trained on small personal sound corpora. We present its architecture, combining a variational DDSP synthesis model, latent smoothing, multi-scale spectral and adversarial losses, and an optional transformer prior, alongside a set of bending operations that intervene directly in the synthesis chain: component limiting, waveshaping, and prior feedback. The Max for Live interface exposes control generation, trajectory sampling, and modulation as primary modes of interaction. Throughout, we thread an affordance analysis arguing that the system's performative character follows from architectural decisions rather than being designed on top of them. The paper contributes both a technical account of the system and a situated affordance analysis of its role in live electronic music performance.
Sources
- Understanding disentangling in $\beta$-VAE
- RAVE: A variational autoencoder for fast and high-quality neural audio synthesis
- Generative timbre spaces: regularizing variational auto-encoders with perceptual metrics
- I'm Sorry for Your Loss: Spectrally-Based Audio Distances Are Bad at Pitch
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment