Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Dongseong Hwang, Prasanth Yadla, Kaan Elgin, Shifas Padinjaru Veettil, Sivanand Achanta, Dipjyoti Paul, Ramya Rasipuram, Tyler Johnson, Emad Soroush, Chung-Cheng Chiu, Zhifeng Chen
cs.SD, cs.CL
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 11 pages, ICASSP
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- MusicLM: Generating Music From Text
- Longformer: The Long-Document Transformer
- Moshi: a speech-text foundation model for real-time dialogue
- Jukebox: A Generative Model for Music
- High Fidelity Neural Audio Compression
- Long-form music generation with latent diffusion
- Fish Audio S2 Technical Report
- Instruction-Following Pruning for Large Language Models
- Efficient Neural Audio Synthesis
- DiffWave: A Versatile Diffusion Model for Audio Synthesis
- AudioGen: Textually Guided Audio Generation
- Generative Spoken Language Modeling from Raw Audio
- MOSNet: Deep Learning based Objective Assessment for Voice Conversion
- Chunked Autoregressive GAN for Conditional Waveform Synthesis
- WaveNet: A Generative Model for Raw Audio
- GLU Variants Improve Transformer
- TS3-Codec: Transformer-Based Simple Streaming Single Codec
- T-Mimi: A Transformer-based Mimi Decoder for Real-Time On-Phone TTS
- Root Mean Square Layer Normalization
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment