Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

arXiv:2601.12966 · cs.SD, cs.CL · Submitted 2026-01-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings".

Tom: A controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for any speaker without requiring explicit Lombard data during training is presented,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So we've seen that this paper is about a controllable text-to-speech system that can synthesize Lombard speech for any speaker without needing explicit Lombard data during training. The core idea is leveraging style embeddings from a large dataset and using principal component analysis to shift those embeddings, which allows them to generate speech at desired Lombard levels.

Jane: Exactly, Tom. The authors are claiming they've developed a method that maps between plain and Lombard speaker embeddings and lets you interpolate between them during inference. This addresses the limitations of existing methods that rely on specific training datasets for each target style.

Lu: It’s neat how they connect the learned style features to specific prosodic attributes like loudness and clarity, which shows a deep understanding of how these characteristics manifest in speech.

Meng: I'm still wondering about the practical details of how they achieve this mapping; knowing *how* it works is important for assessing its real-world viability.

Lalam: The ability to control prosody directly through style embeddings suggests a future where speech synthesis isn't just about reading text, but about precisely crafting the emotional and environmental context of the utterance.

Conclusion: Tom: So, wrapping up on "Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings," we're looking at how the authors managed to synthesize speech for any speaker using style embeddings without ever seeing Lombard data during training, and they showed fine control over the resulting loudness and clarity.

Jane: The implications are that we could have TTS systems that adapt instantly to whatever acoustic environment a user is in, offering much better communication in noisy settings or when speaking to people who have hearing impairments.

Lu: If this technique scales well beyond just English and German prompts as they tested it, the possibilities for multilingual, context-aware speech synthesis become quite expansive.

Meng: From an engineering viewpoint, the ability to control these parameters suggests we could build more flexible assistive technologies that respond dynamically to user needs rather than having fixed settings.

Lalam: This research really shows how AI can move beyond simple text-to-speech and start becoming a powerful tool for tailoring communication precisely to the user's immediate circumstances.

Seymanur Akti, Alexander Waibel

KIT Campus Transfer GmbH (KCT) · Karlsruhe Institute of Technology (KIT) · Carnegie Mellon University (CMU)

cs.SD, cs.CL

Submitted: 2026-01-19

Updated: 2026-10-06

Code: https://github.com/openai/whisper3https:

Project page: https://seymanurakti.github.io/lombard-speech

Importance score: 77/100

The gist: A controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for any speaker without requiring explicit Lombard data during training is presented, which addresses limitations in

Key concepts

Style Embeddings
These are fixed-size numerical representations of prosodic characteristics learned from a large, diverse dataset. They capture the inherent style or manner of speaking—like loudness or clarity—that is independent of the specific speaker's identity. The model uses these embeddings to condition its output speech.
Principal Component Analysis (PCA)
PCA is a mathematical technique used here to analyze style embeddings by finding the most important underlying dimensions (principal components) that capture variance in attributes like loudness and clarity. By projecting the style embedding into this space and shifting specific components, researchers can precisely control desired speech attributes.
FiLM Conditioning
FiLM (Feature-wise Linear Modulation) is a mechanism used to inject style information into the neural network blocks during synthesis. It scales and shifts the model's internal feature outputs based on the extracted style embeddings. This allows the generated speech to be modulated in real-time according to the desired Lombard settings.
Zero-Shot Synthesis
This means synthesizing speech at a specific speaking style (Lombard) using only a standard text input and a general reference audio sample, without needing any separate training data specifically for that style. The system generalizes its learned mapping to produce the target speech directly.

Terminology

Summary

A controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for any speaker without requiring explicit Lombard data during training is presented, which addresses limitations in existing methods that rely on specific training datasets. The core contribution lies in leveraging style embeddings learned from a large, prosodically diverse dataset and manipulating them via Principal Component Analysis (PCA) to generate speech at desired Lombard levels.

The gist: Our method enables zero-shot Lombard speech synthesis by learning a mapping between plain and Lombard speaker embeddings and interpolating between them during inference.

Model Architecture and Baseline

The proposed system is built upon the F5-TTS as the baseline, which is a flow-matching-based TTS model known for its high-quality outputs and built-in duration control. The architecture incorporates components inherited from F5-TTS (blue blocks) alongside new modules (green blocks). To inject speaker information without relying on in-context learning, an ECAPA-TDNN encoder extracts 1024D style features from reference mel-spectrograms. These features are then used via FiLM conditioning to scale and shift the block outputs. Furthermore, to prevent leakage of speaker characteristics, the masked mel-spectrogram inputs are augmented with formant shifts.

Style Embedding Analysis and PCA Manipulation

The research analyzed the learned style embedding space using PCA on both Lombard and articulation datasets to identify correlations with prosodic attributes. Loudness, associated with vocal effort, was analyzed using the AVID corpus, revealing that the first two principal components and measured speech pressure level (SPL), indicating that the style encoder inherently captures loudness variation. Clarity, associated with speaking rate and articulation, was analyzed using the ALBA dataset; PCA showed that the second principal component (PC2) strongly correlates with clarity. The control mechanism involves projecting style embeddings into this PCA space, shifting relevant components according to their variance and user-specified coefficients, and then applying an inverse PCA transformation to obtain manipulated embeddings.

Lombardness Control Mechanism

The system achieves fine-grained control over speech attributes through the manipulation of these style embeddings. The desired Lombard levels are simulated by setting specific parameters: Soft: loudness = -0.5, clarity = -0.5, speed = 1.0, Normal: loudness = 0.0, clarity = 0.0, speed = 1.0, Loud: loudness = 0.5, clarity = 0.5, speed = 0.9, and Very Loud: loudness = 1.0, clarity = 1.0, speed = 0.9. Additionally, the total duration of the output is adjusted by modifying the syllable-to-time ratio using F5-TTS's duration control to enhance perceived clarity during synthesis.

Evaluation and Performance Metrics

Experiments compared F5TTS-Style against F5TTS-Base on English and German prompts, measuring Word Error Rate (WER) for intelligibility, Speaker Similarity (SSIM) for speaker preservation, and UTMOS for perceived naturalness. In the cross-lingual setting with German prompts, F5TTS-Style substantially outperforms the baseline in both intelligibility (lower WER) and naturalness. Robustness to noise was assessed using relative WER (∆WER), which normalizes noisy WER by clean samples. The results demonstrated that control Lombardness enhances noise robustness, with ∆WER dropping sharply at higher Lombardness levels, confirming that intelligibility is preserved even under severe noise conditions (SNR = 1). Speaker identity preservation was confirmed by showing that absolute SSIM remains consistent across all Lombardness levels.

Conclusion and Contribution

The framework successfully enables zero-shot Lombard speech synthesis, generalization to unseen speakers, and direct control over prosody without additional Lombard data. The work introduces F5-TTS Style, a fine-tuned version of F5-TTS Base that supports cross-lingual speaker adaptation and explicit control of Lombard characteristics by leveraging large-scale prosodic style embeddings. This offers a robust solution for controllable Lombard TTS while maintaining speaker identity.

How it works

  1. An ECAPA-TDNN encoder extracts 1024D style features from reference mel-spectrograms to generate the fixed-size style embedding used for conditioning.

  2. The model incorporates these embeddings via FiLM conditioning in the later Diffusion Transformer blocks, where outputs are scaled and shifted by parameters derived from the style embeddings.

  3. Lombardness is controlled by projecting style embeddings into a PCA space, shifting relevant components based on user-specified coefficients, and applying an inverse PCA transformation to obtain manipulated embeddings.

  4. Duration control is applied at inference to adjust speech length based on the desired speaking rate, which is particularly important for simulating clear speech.

  5. The system generates speech conditioned solely on the input text and the style embedding extracted from a reference audio sample, allowing for zero-shot synthesis without Lombard-specific training data.

Improvements for AI systems

Here are specific improvements for AI systems based on the proposed Lombard Speech Synthesis framework, along with what these improved systems can achieve:


) 1. Zero-Shot, Speaker-Agnostic Lombard Speech Synthesis:

The system can synthesize speech at any desired Lombard level (Soft, Normal, Loud, Very Loud) for a speaker whose voice is not present in the training data (zero-shot).

  1. Controllable Prosodic Manipulation:

The system allows explicit modulation of two primary prosodic attributes—loudness and clarity—by shifting specific principal components (PC1 and PC2) of learned style embeddings via PCA manipulation.

  1. Robustness to Acoustic Degradation (Noise Injection):

The improved system can generate speech that maintains high intelligibility even under varying levels of acoustic noise (SNR = 10, 5, 1), demonstrating superior performance compared to baseline models and ground-truth samples in terms of relative Word Error Rate (∆WER).

  1. Speaker Identity Preservation Under Style Modulation:

The system ensures that manipulating Lombard attributes (loudness/clarity) does not degrade the perceived speaker identity, as evidenced by consistent and high Speaker Similarity Scores (SSIM) across all controlled Lombard levels.

  1. Cross-Lingual Adaptability:

By leveraging style embeddings learned from a diverse dataset like Emilia, the system can generalize its control mechanism to generate Lombard speech for speakers in unseen languages (e.g., generating German Lombard speech based on an English speaker's style embedding).

  1. Fine-Grained Duration Control:

The system incorporates duration control at inference, allowing users to adjust the speaking rate (syllable-to-time ratio) to further enhance perceived clarity, which is crucial for simulating clear speech under noise.

This improved AI system can be deployed in applications such as:

  1. Hearing-Assistive Technologies: Providing real-time, high-quality Lombard speech synthesis for hearing aid users in noisy environments without needing personalized Lombard training data.

  2. Robust Speech Recognition Systems (ASR): Enhancing the performance of Automatic Speech Recognition models by synthesizing Lombard versions of speech during inference to simulate noisy conditions, thereby reducing WER in real-world deployment.

  3. Interactive Voice Assistants: Allowing AI assistants to dynamically adjust their output volume and articulation based on perceived environmental noise or the needs of a hearing-impaired interlocutor.

  4. Speech Repair and Error Correction: Generating synthesized speech that compensates for distortions or artifacts in corrupted audio streams, making the resulting speech clearer for downstream processing.

Sources

Related papers