Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
summary
The gist
Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages.
In short
Nord-ParlTTS is an open TTS dataset for Finnish and Swedish derived from wild speech found in Nordic parliament recordings, addressing data scarcity for these languages. The dataset includes 900 hours of Finnish and 5090 hours of Swedish audio, built through a multi-step curation pipeline involving noise reduction, speaker diarization, and quality filtering.
Key concepts
- Dataset Curation Pipeline
- This is the detailed process used to clean and prepare raw audio recordings into usable speech data. For Finnish, it involves standardizing format (24 kHz mono), separating speech from noise using a specific model, identifying different speakers (diarization), and validating transcripts with an ASR model before filtering low-quality samples.
- Speaker Diarization
- This technique automatically identifies and separates different speakers within an audio track. In this project, it is used to isolate individual voices from the parliament recordings, ensuring that the subsequent training data focuses on clear speech segments rather than mixed conversations.
- Objective Evaluation Metrics
- These are quantitative methods used to measure how well a generated TTS model performs. The researchers use metrics like Character Error Rate (CER) when comparing synthetic speech to text, and Cosine Speaker Similarity (SIM) scores to check if the synthesized voice sounds similar to the original prompt speaker.
Terminology used across episodes
This episode discusses
- Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech · Paper Radio
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
The paper
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech · Read on arXiv
Zirui Li, Jens Edlund, Yicheng Gu, Nhan Phan, Lauri Juvela, Mikko Kurimo
Department of Information and Communications Engineering, Aalto University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech".
Tom: Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: We're starting with the title and who put this paper together, Nord-ParlTTS: Finnish and Swedish TTS Dataset from Parliament Speech. Jane, can you explain what the core idea of this dataset is in plain English?
Jane: The main point is that they are providing a resource for text-to-speech development by using speech already existing in the wild—specifically, recordings from Nordic parliaments—to create Finnish and Swedish TTS data. It’s about filling that gap where good speech data is scarce for many languages.
Lu: The authors are a team spanning different institutions, which suggests a strong collaboration between language technology and practical speech processing methods across Europe and Asia.
Meng: I'm looking at the authors—Jens Edlund from KTH Royal Institute of Technology in Sweden, Yicheng Gu from the Chinese University of Hong Kong, and others—and it shows they’ve brought together expertise in both European speech science and large-scale data handling.
Lalam: The diversity of the authors hints at a broad approach to solving this resource problem, combining different methodological strengths to build a comprehensive solution for these languages.
The paper's summary: Tom: So, let's talk about what the paper actually says about how they built this Nord-ParlTTS dataset. Jane, walk us through the main steps they took in turning those raw parliament recordings into usable TTS data.
Jane: They adapted an existing pipeline called the Emilia data processing pipeline for both languages. For Finnish, they standardized the audio to a twenty-four kHz mono-channel format and used a pretrained UVR-MDX-Net Inst thirty-four model to separate the speech from noise.
Lu: That's interesting because using a specialized noise separation model suggests they were trying to handle the inherent acoustic variability of live parliamentary recordings before moving on to speaker diarization.
Meng: And then they used speaker diarization from Pyannote and a fine-grained Voice Activity Detection model, Silero VAD, which is crucial for accurately isolating individual speaking turns in complex meetings.
Lalam: Isolating those turns is key because it ensures that the data we feed into the TTS model represents distinct speakers rather than just a continuous stream of noise and speech.
The paper's improvements: Tom: Beyond just collecting the data, what improvements did the authors suggest in how this dataset should be used or how models should be trained? Jane, what are the key methodological tweaks they propose to make this data better for actual TTS training?
Jane: They point out that simply using an ASR corpus derived from these recordings isn't enough because segmenting based on word boundaries makes it hard for TTS models to learn correct pronunciations. So, they suggest a different approach to preparing the transcripts.
Lu: They specifically mention correcting transcripts for hesitations and replacing spontaneous language with written equivalents for clarity, which is a necessary step before training any model on that text.
Meng: From a practical standpoint, having those corrections done beforehand saves immense time in the downstream model development phase because it removes some of the ambiguity in the text input.
Lalam: It shows they aren't just handing over raw audio; they are thinking about how to preprocess the language data itself to maximize its utility for high-quality synthesis.
Conclusion: Tom: Alright, we've covered a lot about Nord-ParlTTS today, looking at the dataset creation and the suggested improvements. Jane, could you give us your final thoughts on what this means for TTS development?
Jane: It really means that by providing this open dataset of nine hundred Finnish hours and five thousand ninety Swedish hours, we are seriously narrowing the resource gap between high-resource languages like English and these Nordic ones <ref:2509.17988#pg0>. It makes high-quality synthesis more accessible.
Lu: I think the real potential here is in how this data fuels cross-lingual transfer learning; if we can use these unified evaluation sets, we could quickly bootstrap systems for new Nordic languages by leveraging the structures already learned from Finnish and Swedish.
Meng: For me, the practical impact is that small teams or independent developers won't have to start from scratch collecting massive amounts of audio; they can build competitive TTS systems much faster with this data foundation.
Lalam: I really see this advancing culture because it democratizes access to high-fidelity voice synthesis, meaning more diverse voices can be created and used in applications across the region.
Tom: Fantastic points everyone. So, to wrap up on Nord-ParlTTS: it’s a significant contribution by providing large-scale data from parliament speech for Finnish and Swedish TTS, using an adapted pipeline that addresses transcript quality issues, with the goal of closing the resource gap in these languages. That's all for this segment.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck