SCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis
cs.SD, cs.AI, cs.CR
Submitted: 2025-08-11
Updated: 2025-08-11
Journal ref: 2025 International Conference of the Biometrics Special Interest Group (BIOSIG)
DOI: 10.1109/BIOSIG65492.2025.11358014
Code: https://github.com/SWivid/F5-TTS
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain underexplored in the speech domain.
Terminology
Abstract
Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain underexplored in the speech domain. To address this gap, we introduce the Speaker Characteristics Deepfake (SCDF) dataset: a novel, richly annotated resource enabling systematic evaluation of demographic biases in deepfake speech detection. SCDF contains over 237,000 utterances in a balanced representation of both male and female speakers spanning five languages and a wide age range. We evaluate several state-of-the-art detectors and show that speaker characteristics significantly influence detection performance, revealing disparities across sex, language, age, and synthesizer type. These findings highlight the need for bias-aware development and provide a foundation for building non-discriminatory deepfake detection systems aligned with ethical and regulatory standards.
Sources
- Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and Evaluations
- XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
- OpenVoice: Versatile Instant Voice Cloning
- XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment