P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution
eess.AS, cs.LG, cs.SD
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: Under review
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models
- Audio Super Resolution using Neural Networks
- CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space
- A2SB: Audio-to-Audio Schrodinger Bridges
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions