Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models
cs.CR, cs.LG, cs.SD, eess.AS
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/resemble-ai/chatterbox
Terminology
Sources
- SLMIA-SR: Speaker-Level Membership Inference Attacks against Speaker Recognition Systems
- Towards Black-Box Membership Inference Attack for Diffusion Models
- Detecting Pretraining Data from Large Language Models
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
- A Survey on Neural Speech Synthesis
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
- Membership Inference Attacks Against Text-to-image Generation Models
- Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
- Do Membership Inference Attacks Work on Large Language Models?
- Voice Cloning: Comprehensive Survey
- Qwen3-TTS Technical Report
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy
- OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
- Quantifying Memorization Across Neural Language Models
- Content-Dependent Fine-Grained Speaker Embedding for Zero-Shot Speaker Adaptation in Text-to-Speech Synthesis
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
- Qwen3-Omni Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs