Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS
eess.AS, cs.CL, cs.CV
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/IIP-Sogang/olkavs-avspeech
Terminology
Sources
- Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language
- LoRA: Low-Rank Adaptation of Large Language Models
- The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions