A survey of AI-generated voices and their detection
cs.AI, cs.CV
Submitted: 2026-08-15
Updated: 2026-08-15
Terminology
Sources
- Listening for Expert Identified Linguistic Features: Assessment of Audio Deepfake Discernment among Undergraduate Students
- XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
- VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
- DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice Conversion
- PointGAT: A quantum chemical property prediction model integrating graph attention and 3D geometry
- Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network
- MaskCycleGAN-VC: Learning Non-parallel Voice Conversion with Filling in Frames
- Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
- SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
- Towards Robust FastSpeech 2 by Modelling Residual Multimodality
- BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
- StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion
- CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
- MOSNet: Deep Learning based Objective Assessment for Voice Conversion
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
- Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation
- WaveFake: A Data Set to Facilitate Audio Deepfake Detection
- Passive Deepfake Detection Across Multi-modalities: A Comprehensive Survey
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection