CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

arXiv:2608.11590 · cs.SD, cs.LG · Submitted 2026-08-13 · Read on arXiv

Haowei Lou, Hye-Young Paik, Dai Jia, Kai Li, Lina Yao

UNSW Sydney · Dolby Laboratories

cs.SD, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Project page: https://haoweilou.github.io/CookVoice

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation Summary This paper proposes CookVoice, a unified non-autoregressive (NAR) framework for multimodal, multi-style,

Terminology

Summary

CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

Summary

This paper proposes CookVoice, a unified non-autoregressive (NAR) framework for multimodal, multi-style, and multi-task human voice generation. The framework decomposes human voice into three key factors—content, prosody, and style—enabling both speech and singing voice generation within a single model. CookVoice supports a wide range of tasks including text-to-speech (TTS), text-to-singing voice (TTSV), style-controllable generation, voice mimicry, voice conversion, and voice editing.

The core design of CookVoice is a flexible frame-level alignment strategy that maps text, style, and prosody control signals onto the frame-level of the spectrogram. This allows explicit and fine-grained temporal control, unlike autoregressive (AR) models that generate acoustic tokens sequentially and model duration and prosody implicitly, or NAR systems that rely on restricted alignments such as phoneme-level durations or note-to-phoneme mapping. CookVoice explicitly expands different control signals to the acoustic frame level and allows them to be combined freely.

The framework formulates human voice generation as conditional latent acoustic generation. Content is represented by phoneme sequences derived from text or lyrics. Prosody can be controlled through discrete signals (lexical tones, stresses, MIDI notes) or continuous frame-level F0 contours extracted from a reference voice. Style can be provided by either text-based descriptions or reference voice prompts. The temporally aligned representations are integrated through a multimodal adaptive fusion module and passed to a flow-matching Diffusion Transformer (DiT) for high-quality latent acoustic generation.

CookVoice uses a latent acoustic encoding scheme based on a HiFi-GAN-style autoencoder. The style encoder generates style embeddings from either text (using a frozen MPNet sentence encoder) or reference voice (using a trainable Transformer encoder with attentive pooling). The content encoder uses a multilingual Grapheme-to-Phoneme (G2P) module to convert text into phonemes and prosody tokens, with duration expansion to match the temporal dimension of the latent acoustic embedding. The prosody encoder handles both discrete signals (lexical prosody tokens for speech, musical note tokens for singing) and continuous signals (F0 contours extracted from reference voice waveforms).

During training, CookVoice uses a condition-switching strategy where style and prosody conditions are randomly sampled from different sources (text vs. voice style, discrete vs. continuous prosody) at the sample level within each batch. This allows a single model to support multiple tasks without changing architecture or training objectives.

The model is trained on approximately 123k voice samples (168 hours) from 6,361 speakers, combining multiple open-source bilingual (English and Chinese) speech and singing voice datasets. The model uses a DiT-S backbone with 43.51 million parameters and is trained on a single NVIDIA RTX 5090 GPU for 800K steps.

Experimental results show that CookVoice achieves generation quality comparable to existing TTS and TTSV baselines while providing stronger style and prosody controllability. For TTS with text-based style and discrete prosody control, CookVoice improves Style Similarity (S-SIM) by 41.48% and F0-CORR by 121.10% compared to the best baselines. For TTSV with voice-based style and continuous F0 control, it improves S-SIM by 7.84% and F0-CORR by 18.25%. CookVoice achieves S-SIM of 91.65% for TTS and 95.00% for TTSV, with F0-CORR of 0.7102 and 0.8425 respectively.

In terms of perceptual quality, CookVoice achieves a best MOS of 3.98 for TTS (close to ground truth of 4.05) and 3.40 for TTSV (comparable to Vevo2's 3.42). It obtains the highest MC-MOS score (0.28) among compared singing voice systems, indicating better perceived melody consistency.

The paper also analyzes the effects of different control signals. Voice-based style conditioning consistently outperforms text-based conditioning across all metrics for TTS, with the largest difference in S-SIM (0.932 vs. 0.199 normalized scores). Continuous prosody signals (F0) provide clear advantages over discrete signals in prosody-related metrics, with CONT achieving scores of 0.844 for F0-RMSE and 0.892 for F0-CORR compared to 0.230 and 0.221 for DISC in TTS.

Regarding inference efficiency, CookVoice achieves real-time generation with an RTF of 0.04 using as few as 4 ODE steps. The paper recommends 4–8 ODE steps as the optimal operating range, where style similarity and prosody fidelity stabilize while intelligibility metrics achieve their best trade-off. Compared to Vevo2 (the most directly comparable baseline supporting both speech and singing), CookVoice uses only 4.99% of the parameters, 20.03% of the CUDA memory, and 0.27% of the RTF.

The main contributions are: (1) proposing CookVoice as a unified framework for multi-task human voice generation by decomposing voice into style, content, and prosody; (2) designing a flexible frame-level alignment mechanism for fine-grained controllability supporting multimodal style and prosody control; (3) demonstrating higher style and prosody controllability than existing baselines while maintaining efficient inference and competitive generation quality with a lightweight architecture.

Limitations include that CookVoice has not been scaled up (only uses DiT-S with 43.51M parameters trained on 168 hours of data), and its potential applicability to broader audio generation domains (music, instrumental sound, general audio) has not been investigated.

Improvements for AI systems

Improvements to AI Systems:

  1. Unified Multi-Modal Voice Generation with Frame-Level Control
  • Implement a single non-autoregressive model that handles TTS, text-to-singing, voice conversion, voice editing, and style mimicry without task-specific fine-tuning.

  • Use explicit frame-level alignment of content, prosody, and style signals to spectrogram frames, enabling fine-grained temporal control (e.g., adjusting pitch or emotion at specific syllables or notes) that autoregressive models cannot achieve.

  1. Multimodal Style and Prosody Fusion
  • Integrate a multimodal adaptive fusion module that combines text-based style descriptions (via frozen sentence encoders) and reference voice embeddings (via trainable transformer encoders) with discrete prosody tokens (lexical tones, MIDI notes) and continuous F0 contours.

  • Allow dynamic switching between these condition sources during inference, enabling users to control style via text, voice prompt, or both simultaneously.

  1. High-Fidelity Latent Acoustic Generation with Flow-Matching DiT
  • Replace autoregressive token prediction with a flow-matching Diffusion Transformer (DiT) that generates latent acoustic embeddings directly, improving prosody fidelity (F0-CORR up to 0.8425) and style similarity (S-SIM up to 95%) while reducing inference latency (RTF 0.04 with 4 ODE steps).
  1. Condition-Switching Training for Multi-Task Generalization
  • Train the model with random per-sample condition switching (text vs. voice style, discrete vs. continuous prosody) within each batch, allowing a single architecture to generalize across tasks without architectural changes—enabling zero-shot style transfer and voice mimicry from unseen references.
  1. Lightweight and Efficient Real-Time Deployment
  • Use a compact DiT-S backbone (43.51M parameters) trained on 168 hours of data, achieving real-time generation on consumer hardware (e.g., single RTX 5090) with 4–8 ODE steps, reducing parameter count by 95% and memory by 80% compared to baselines like Vevo2.
  1. Improved Controllability Metrics
  • Provide explicit control over prosody via both discrete (lexical tones, MIDI) and continuous (F0) signals, with continuous signals showing 3.7x better F0-CORR and 3.6x better F0-RMSE than discrete—enabling precise pitch contour editing for expressive speech or singing.

What the Improved AI System Can Do:

  • Generate speech or singing from text or lyrics with user-specified style (e.g., angry whisper or operatic soprano) and fine-grained prosody (e.g., rising pitch on a specific word or note).

  • Mimic a reference voice (speech or singing) in real-time, preserving style and F0 contours while changing content.

  • Edit existing audio by modifying prosody or style at frame-level precision (e.g., change emotion of a single phrase without re-synthesizing the whole utterance).

  • Perform voice conversion across speakers and languages (English/Chinese) with high style similarity (91.65% for TTS, 95% for TTSV).

  • Run on edge devices or real-time applications (e.g., virtual assistants, dubbing, singing avatars) due to low latency (RTF 0.04) and small memory footprint.

  • Support multimodal control—users can mix text descriptions and voice prompts, or switch between discrete and continuous prosody inputs, for flexible creative expression.

Abstract

Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for specific tasks and often rely on task-dependent architectures, control signals, or autoregressive decoding, limiting fine-grained controllability and inference efficiency. In this paper, we propose CookVoice, a unified framework for multimodal, multi-style, and multi-task human voice generation. CookVoice decomposes the human voice into three key factors: content, prosody, and style, enabling both speech and singing voice generation within a unified model. To achieve precise and flexible controllability, we design a flexible alignment strategy that maps text, style, and prosody control signals onto the frame-level of spectrogram. This design allows CookVoice to support a wide range of tasks, including text-to-speech, text-to-singing voice, style-controllable generation, voice mimicry, voice conversion, and voice editing. Experimental results show that CookVoice achieves generation quality comparable to existing Text-to-Speech and text-to-singing voice baselines, while providing stronger style and prosody controllability. Moreover, CookVoice achieves comparable performance to large-scale baselines with only 43.51 million parameters and efficient inference using as few as 4 ODE steps, making it a practical solution for real-world human voice generation applications. Demo page is available at https://haoweilou.github.io/CookVoice/.

Sources

Related papers