Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning
cs.SD, cs.CL, eess.AS
Submitted: 2026-09-24
Updated: 2026-09-24
Project page: https://yoomee-cho.github.io/accent-analogy-guidance
Terminology
Sources
- OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
- Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
- X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
- THCHS-30 : A Free Chinese Speech Corpus
- Classifier-Free Diffusion Guidance
- MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
- Selective Classifier-free Guidance for Zero-shot Text-to-speech
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS
- Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech
- CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations
- Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment