GEPARD - Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue

arXiv:2609.04222 · eess.AS, cs.CL, cs.LG, cs.SD · Submitted 2026-07-04 · Read on arXiv

eess.AS, cs.CL, cs.LG, cs.SD

Submitted: 2026-07-04

Updated: 2026-07-04

Comments: Technical Report. 37 pages, 11 figures, Demo samples, code, and open weights: https://huggingface.co/nineninesix/gepard-1.0 ; https://github.com/nineninesix-ai/gepard-inference . Affiliation: Nineninesix, Inc

Code: https://github.com/nineninesix-ai/gepard-inference

License: http://creativecommons.org/licenses/by/4.0/

The gist: We present GEPARD (Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue), a streaming text-to-speech model for real-time spoken dialogue.

Terminology

Abstract

We present GEPARD (Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue), a streaming text-to-speech model for real-time spoken dialogue. GEPARD generates speech autoregressively with an LLM backbone - text and audio embeddings are trained together in a single decoder-only model - and decodes it to a waveform with an FSQ-based neural codec, streaming audio chunk-by-chunk as text arrives. Our central goal is a TTS architecture served by a standard LLM engine (vLLM) without modifying its compute kernels. This defines the overarching design principle: the backbone is a standard full-attention transformer, while all non-trivial auxiliary mechanisms - zero-shot voice cloning, text augmentation, and classifier-free guidance - are moved out of the autoregressive decode loop into prefill, or distilled directly into the weights. On streaming end-to-end inference, a single stream reaches a Real-Time Factor of about 0.067 (roughly 15x faster than real-time); under 256 concurrent streams the system reaches an aggregate speedup of about 204x on a single server-class GPU. We detail: (1) system-level solutions for vLLM-native serving; (2) the "short register" (1-2 word) failure mode of autoregressive speech decoders, with diagnostic probes and a mitigation; and (3) distillation of two-pass classifier-free guidance over text into single-pass weights via Direct Preference Optimization (DPO).

Related papers