Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech
cs.SD, cs.AI
Submitted: 2026-09-23
Updated: 2026-09-23
Terminology
Sources
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
- Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
- Continual Speaker Identity Unlearning with Minimal Interference
- EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
- Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
- GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech
- Adam: A Method for Stochastic Optimization
- LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
- FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment