Provable Speech Attributes Conversion via Latent Independence
cs.SD, cs.AI
Submitted: 2025-10-06
Updated: 2026-09-24
Project page: https://jsvir.github.io/ivc
Terminology
Sources
- Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
- Sparse Binarization for Fast Keyword Spotting
- Speech Resynthesis from Discrete Disentangled Self-Supervised Representations
- Voice Conversion With Just Nearest Neighbors
- emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
- ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
- Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
- Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
- GenVC: Self-Supervised Zero-Shot Voice Conversion
- RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
- Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data
- Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training
- SpeechBrain: A General-Purpose Speech Toolkit
- FineGates: LLMs Finetuning with Compression using Stochastic Gates
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment