Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech".
Tom: Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: We're starting with the title and who put this paper together, Nord-ParlTTS: Finnish and Swedish TTS Dataset from Parliament Speech. Jane, can you explain what the core idea of this dataset is in plain English?
Jane: The main point is that they are providing a resource for text-to-speech development by using speech already existing in the wild—specifically, recordings from Nordic parliaments—to create Finnish and Swedish TTS data. It’s about filling that gap where good speech data is scarce for many languages.
Lu: The authors are a team spanning different institutions, which suggests a strong collaboration between language technology and practical speech processing methods across Europe and Asia.
Meng: I'm looking at the authors—Jens Edlund from KTH Royal Institute of Technology in Sweden, Yicheng Gu from the Chinese University of Hong Kong, and others—and it shows they’ve brought together expertise in both European speech science and large-scale data handling.
Lalam: The diversity of the authors hints at a broad approach to solving this resource problem, combining different methodological strengths to build a comprehensive solution for these languages.
The paper's summary: Tom: So, let's talk about what the paper actually says about how they built this Nord-ParlTTS dataset. Jane, walk us through the main steps they took in turning those raw parliament recordings into usable TTS data.
Jane: They adapted an existing pipeline called the Emilia data processing pipeline for both languages. For Finnish, they standardized the audio to a twenty-four kHz mono-channel format and used a pretrained UVR-MDX-Net Inst thirty-four model to separate the speech from noise.
Lu: That's interesting because using a specialized noise separation model suggests they were trying to handle the inherent acoustic variability of live parliamentary recordings before moving on to speaker diarization.
Meng: And then they used speaker diarization from Pyannote and a fine-grained Voice Activity Detection model, Silero VAD, which is crucial for accurately isolating individual speaking turns in complex meetings.
Lalam: Isolating those turns is key because it ensures that the data we feed into the TTS model represents distinct speakers rather than just a continuous stream of noise and speech.
The paper's improvements: Tom: Beyond just collecting the data, what improvements did the authors suggest in how this dataset should be used or how models should be trained? Jane, what are the key methodological tweaks they propose to make this data better for actual TTS training?
Jane: They point out that simply using an ASR corpus derived from these recordings isn't enough because segmenting based on word boundaries makes it hard for TTS models to learn correct pronunciations. So, they suggest a different approach to preparing the transcripts.
Lu: They specifically mention correcting transcripts for hesitations and replacing spontaneous language with written equivalents for clarity, which is a necessary step before training any model on that text.
Meng: From a practical standpoint, having those corrections done beforehand saves immense time in the downstream model development phase because it removes some of the ambiguity in the text input.
Lalam: It shows they aren't just handing over raw audio; they are thinking about how to preprocess the language data itself to maximize its utility for high-quality synthesis.
Conclusion: Tom: Alright, we've covered a lot about Nord-ParlTTS today, looking at the dataset creation and the suggested improvements. Jane, could you give us your final thoughts on what this means for TTS development?
Jane: It really means that by providing this open dataset of nine hundred Finnish hours and five thousand ninety Swedish hours, we are seriously narrowing the resource gap between high-resource languages like English and these Nordic ones <ref:2509.17988#pg0>. It makes high-quality synthesis more accessible.
Lu: I think the real potential here is in how this data fuels cross-lingual transfer learning; if we can use these unified evaluation sets, we could quickly bootstrap systems for new Nordic languages by leveraging the structures already learned from Finnish and Swedish.
Meng: For me, the practical impact is that small teams or independent developers won't have to start from scratch collecting massive amounts of audio; they can build competitive TTS systems much faster with this data foundation.
Lalam: I really see this advancing culture because it democratizes access to high-fidelity voice synthesis, meaning more diverse voices can be created and used in applications across the region.
Tom: Fantastic points everyone. So, to wrap up on Nord-ParlTTS: it’s a significant contribution by providing large-scale data from parliament speech for Finnish and Swedish TTS, using an adapted pipeline that addresses transcript quality issues, with the goal of closing the resource gap in these languages. That's all for this segment.
Zirui Li, Jens Edlund, Yicheng Gu, Nhan Phan, Lauri Juvela, Mikko Kurimo
Department of Information and Communications Engineering, Aalto University
eess.AS, cs.CL
Submitted: 2025-09-22
Updated: 2026-02-07
Comments: Accepted by ICASSP 2026. 5 pages, 2 figures
Code: https://github.com/TRvlvr/model
Project page: https://gryffindorli.github.io/nord-parl-tts-demo
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 77/100
The gist: Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages.
Key concepts
- Dataset Curation Pipeline
- This is the detailed process used to clean and prepare raw audio recordings into usable speech data. For Finnish, it involves standardizing format (24 kHz mono), separating speech from noise using a specific model, identifying different speakers (diarization), and validating transcripts with an ASR model before filtering low-quality samples.
- Speaker Diarization
- This technique automatically identifies and separates different speakers within an audio track. In this project, it is used to isolate individual voices from the parliament recordings, ensuring that the subsequent training data focuses on clear speech segments rather than mixed conversations.
- Objective Evaluation Metrics
- These are quantitative methods used to measure how well a generated TTS model performs. The researchers use metrics like Character Error Rate (CER) when comparing synthetic speech to text, and Cosine Speaker Similarity (SIM) scores to check if the synthesized voice sounds similar to the original prompt speaker.
Terminology
Summary
Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages. Nord-ParlTTS presents an open TTS dataset for Finnish and Swedish based on speech found in the wild, addressing this resource gap between high- and lowerresourced languages.
The gist: Nord-ParlTTS is an open TTS dataset for Finnish and Swedish based on speech found in the wild, extracted from Nordic parliament recordings, resulting in 900 hours of Finnish and 5090 hours of Swedish TTS data.
Dataset Curation
The dataset is built using an adapted version of the Emilia data processing pipeline.
For Finnish, the process involves several steps: first, extracting the audio track and standardizing it to a 24 kHz mono-channel format
; then, separating it from noise using a pretrained UVR-MDX-Net Inst 34 model
and performing speaker diarization using speaker-diarization-3.1 from Pyannote [15].
A fine-grained VAD is conducted on the diarized speech using Silero VAD [16].
To validate the transcript, an extra ASR model, specifically a Wav2Vec2-large model pretrained and fine-tuned solely on Finnish [17],
is used to check the Whisper model’s transcript. The pipeline concludes with DNSMOS [18] prediction,
where samples with a DNSMOS P.835 OVRL score less than 3.0 to be excluded.
For Swedish, the processing utilizes the RixVox-v2 data, involving resampling to 24 kHz
and using the same speech separation model for noise cleaning. Utterances are cut based on existing sentence-level timestamps if gaps exceed 2 seconds,
and speaker diarization is used to exclude utterances with multiple speakers. Intelligibility is evaluated by transcribing the utterance using Swedish Whisper-large [19]
and filtering out those with a Word Error Rate (WER) larger than 10%. Finally, utterances are filtered based on the DNSMOS score, excluding those with a score less than 3.0,
yielding the final 5090-hour Swedish dataset.
Evaluation Sets and Benchmarking
To support model development, the paper proposes unified evaluation sets following [1, 10].
For Finnish, they curate 500 prompt-target pairs from Perso Synteesi [11],
ensuring a gender balance by sampling 250 prompts from male speakers and 250 prompts from female speakers.
All prompt and target speeches are constrained to be between 3 seconds and 20 seconds
with content having more than 10 characters.
For Swedish, the evaluation dataset is curated from CommonVoice [20] Swedish 22.0, keeping utterances between 3 to 20 seconds with non-empty gender labels.
They sample a maximum of 30 utterances to create the pool of candidates
per speaker, balancing the number of utterances per gender and per speaker. A native Swedish speaker then performs a bespoke curation tool to filter out speech with unclear speech, low volume, strong non-native accents, and background noise,
before forming 500 prompt-target pairs.
Model Training and Objective Evaluation
The researchers train two non-autoregressive (NAR) TTS models: Matcha-TTS [21] and F5-TTS [1], both diffusion-based. For Matcha-TTS, they use monotonic alignment search (MAS) [22]
to find optimal alignment, replacing the speaker embedding table with a pretrained SimAMResNet Speaker Encoder from WeSpeaker [25],
and strengthening it with Classifierfree Guidance (CFG) [26].
The Finnish model uses character sequences directly due to its near one-to-one phoneme-to-grapheme mapping.
For F5-TTS, they use a stack of Diffusion Transformers (DiT) for implicit alignment. Both models are trained on 1 Nvidia A100 GPU for 500k updates with a batch size of 64.
The Swedish model is trained using Phonemizer [27] for Swedish G2P.
Objective evaluation involves synthesizing content and using ASR models to calculate the Character Error Rate (CER) between synthetic speech and ground truth content. They also calculate the Cosine Speaker Similarity (SIM) between synthetic speech and prompt speech, utilizing Wav2Vec2-large [17] for Finnish ASR and Swedish Whisper-large [19] for Swedish ASR, along with a WavLM-large-based speaker verification model for SIM.
Subjective Evaluation
For subjective evaluation, they invite participants recruited via Prolific7.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the Nord-Parl-TTS dataset and its associated research, along with what those improved systems could achieve:
The core contribution is a high-quality, in-the-wild TTS dataset for Finnish and Swedish derived from parliamentary speech. Improvements can focus on leveraging this data for cross-lingual generalization, robustness against noise/variability, and fine-grained control over prosody.
Here are the specific improvements:
-
-
Extract highly robust, speaker-invariant representations from the Nord-Parl-TTS data (e.g., using the DNSMOS filtering and diarization techniques) to train more generalized TTS models that are less sensitive to acoustic noise or spontaneous speech variations encountered in real-world parliamentary recordings.
-
-
Implement cross-lingual transfer learning mechanisms, specifically utilizing the unified evaluation sets for Finnish and Swedish TTS (as proposed in Section 1), to enable a single model trained on one language (e.g., Finnish) to achieve strong performance on the other (Swedish) by leveraging shared linguistic structures learned during pre-training.
-
-
Integrate explicit prosodic control mechanisms derived from the analysis of duration and DNSMOS scores in the dataset to allow TTS systems to generate speech with more natural, contextually appropriate rhythm and emphasis, moving beyond simple text-to-speech towards expressive speech generation suitable for narration or synthetic dialogue.
-
-
Leverage the speaker similarity metric (SMOS) from the evaluation phase to develop a speaker embedding system that can accurately clone or synthesize speech in a target language while preserving the unique vocal characteristics of a specific speaker, even when that speaker's voice is spoken in a parliamentary context.
-
-
Develop specialized pre-processing pipelines (like the one detailed in Section 3) for handling low-resource languages by adapting established high-resource processing methods (e.g., adapting the Emilia pipeline modifications to other Nordic languages like Danish or Norwegian), thereby drastically lowering the barrier to entry for creating high-quality datasets for previously underserved linguistic groups.
-
-
Optimize the diffusion model architectures (Matcha-TTS and F5-TTS) by incorporating explicit attention mechanisms that better capture sentence-level prosody, ensuring that the generated speech adheres strictly to the natural timing and phrasing observed in parliamentary discourse, leading to higher intelligibility metrics (lower CER).
The improved AI systems can achieve the following:
-
-
The resulting TTS systems will generate Finnish and Swedish speech with significantly reduced errors (CER improvement) and better speaker fidelity compared to models trained on studio data alone, making them viable for real-time applications in news synthesis or digital assistants.
-
-
A cross-lingual system will allow researchers to rapidly prototype high-quality TTS systems for new Nordic languages by transferring knowledge from established datasets (Finnish/Swedish) without needing to collect massive amounts of new, labeled data for that specific language immediately.
-
-
The system can generate speech with nuanced emotional or narrative delivery, as the prosody modeling ensures the rhythm and emphasis align with natural human speaking patterns observed in complex discourse like parliamentary debate.
-
A high-fidelity voice cloning tool will be able to synthesize a target speaker's voice in either Finnish or Swedish with superior speaker similarity, enabling more realistic synthetic media creation for audiobooks or virtual assistants.
-
-
The AI infrastructure will become significantly more democratized; researchers and small teams can build competitive TTS systems for low-resource languages by simply adapting the proven data processing pipeline, effectively closing the resource gap in speech synthesis accessibility across the Nordic region.
-
-
The final synthesized speech output will exhibit superior temporal accuracy and natural flow, making it suitable for high-stakes applications where precise timing and intelligibility are critical, such as automated transcription systems or advanced voice interfaces.
Abstract
Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages. We present Nord-Parl-TTS, an open TTS dataset for Finnish and Swedish based on speech found in the wild. Using recordings of Nordic parliamentary proceedings, we extract 900 hours of Finnish and 5090 hours of Swedish speech suitable for TTS training. The dataset is built using an adapted version of the Emilia data processing pipeline and includes unified evaluation sets to support model development and benchmarking. By offering open, large-scale data for Finnish and Swedish, Nord-Parl-TTS narrows the resource gap in TTS between high- and lower-resourced languages.
Sources
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions