Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean
cs.CL
Submitted: 2026-06-25
Updated: 2026-09-24
Terminology
Sources
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
- VoxCPM2 Technical Report
- Scaling Speech Technology to 1,000+ Languages
- Joint Khmer Word Segmentation and Part-of-Speech Tagging Using Deep Learning
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
- Exploring Efficient-tuning Methods in Self-supervised Speech Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering