myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR
Ye Kyaw Thu, Ye Bhone Lin, Thura Aung, Htet Arkar, Myat Oo Swe, Thet Htet San, Min Thiha Tun, Thazin Myint Oo, Thepchai Supnithi
National Electronics and Computer Technology Center · Language Understanding Laboratory · King Mongkut's University of Technology Thonburi · King Mongkut's Institute of Technology Ladkrabang
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/iver56/audiomentations
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 81/100
The gist: This paper presents myMediWhisper, a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by native speakers.
Terminology
Summary
This paper presents myMediWhisper, a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by native speakers. The authors fine-tune Whisper models using full fine-tuning (FFT) and parameter-efficient fine-tuning (PEFT) with LoRA. To evaluate robustness, they apply waveform- and spectrogram-level data augmentation under controlled noise and simulated room acoustics. While augmentation reduces performance on clean speech, it significantly improves robustness in noisy and reverberant environments across FFT and PEFT settings. The best-performing system, fully fine-tuned myMediWhisper-Medium without augmentation, achieves a state-of-the-art Word Error Rate (WER) of 23.44%, outperforming much larger general-domain fine-tuned models.
The paper addresses three research questions: RQ1 compares FFT and PEFT fine-tuning strategies for Burmese medical ASR; RQ2 examines how data augmentation influences recognition accuracy and robustness; RQ3 evaluates how well models generalize to noise and room acoustic variations. The contributions include a publicly available Burmese medical speech corpus, Whisper-based ASR benchmarks using FFT and PEFT, and a systematic analysis of robustness under simulated noise and room acoustic variations.
For dataset preparation, the authors adopted Burmese translations of the Samson Handbook of PLAB 2 and Clinical Assessment, with all translated sentences manually verified by two native Burmese speakers. The speech corpus was collected from nine native Burmese speakers (two male and seven female), recorded using built-in microphones on personal devices at 16 kHz. After quality verification, the corpus contains 28 hours, 6 minutes, and 36 seconds of high-quality Burmese medical speech. From the verified corpus, 52.95 hours were allocated as training data, with 2.87 hours reserved for evaluation. Room acoustics were simulated for each training audio sample using the Pyroomacoustics library, generating impulse responses from L-shaped and random-distance room configurations.
Data augmentation involved randomly sampling 50% of the original training set and applying three waveform-level augmentation methods (random time shifting, pitch shifting, and additive Gaussian noise), followed by two spectrogram-level methods (time masking and frequency masking). This expanded the train data from 52.95 hours to 198.63 hours, a 3.75× increase. Hyperparameters included a batch size of 4, one epoch, and a learning rate of 1×10−5, with experiments conducted on dual NVIDIA Tesla T4 GPUs.
Results show that external baselines struggle in this specialized domain: facebook/MMS-1B yields a WER of 37.90%, and whisper-large-v3-myanmar achieves 32.03%. Domain adaptation leads to massive accuracy gains over vanilla zero-shot Whisper variants, with FFT yielding relative WER reductions of up to 94.73% for Whisper-Base and 90.85% for Whisper-Medium without augmentation. PEFT via Rank-Stabilized LoRA (r=128, α=256) enables stable adaptation of Whisper-Large-v2, achieving a WER reduction of 72.13% without augmentation (WER=41.57%) and 68.07% with augmentation (WER=47.63%). PEFT models exhibit higher real-time factors (RTF) than FFT counterparts, indicating inference latency penalties.
Robustness evaluation under varying SNR conditions shows that data augmentation generally improves robustness at lower SNR levels for medium and larger models, while smaller models benefit less consistently. Room acoustic testing with three room types (short distance, long distance, L-shape) shows WER increases under long-distance and L-shaped configurations, with limited gender bias across room types.
Error analysis of the most robust model (myMediWhisper Medium Aug) identified three main error categories: phonological and orthographic substitutions (e.g., /bÉ/ → /pÉ/ with n=36, reflecting acoustic ambiguity between voiced and unvoiced bilabial plosives), function-word deletions (the unstressed nominalizing prefix /P@/ was the most omitted token with n=400, along with polite markers /bà/ n=254, object markers /ěò/ n=220, and realis verb markers /dÈ/ n=202), and weak-prefix insertions (spurious prefix additions like /P@/ n=162 and postpositional markers such as /ka/ n=74).
The paper concludes that FFT delivers the highest precision when hardware permits, yielding the state-of-the-art WER of 23.44% with myMediWhisper-Medium. PEFT via high-rank rsLoRA serves as a vital scaling mechanism for adapting Whisper-Large-v2 under memory constraints, albeit with increased RTF latency. Multi-level data augmentation highlights a critical accuracy–robustness trade-off, with acoustic variations introducing minor penalties on clean test sets but substantially safeguarding performance in noisy and reverberant environments. Residual model failures are predominantly driven by acoustic voicing overlaps and deletion of weak nominal prefixes and grammatical particles due to low acoustic energy in continuous speech.
Limitations include reliance on simulated noise profiles and synthetic room impulse responses, which may not fully capture real-world clinical dynamics such as multi-speaker babble, non-stationary device noise, and microphone distortions. The dataset remains constrained in acoustic scale, speaker diversity, and regional dialect representation. Future work will scale the corpus with additional clinical recordings, expand beyond ASR toward speech translation and medical named entity recognition, and explore alternative architectures incorporating structural constraints targeting Burmese tonal pitch contours and phonetic variations in localized medical loanwords.
Improvements for AI systems
Improvements to AI Systems:
- Domain-Adaptive Fine-Tuning Pipeline for Low-Resource Medical ASR
-
Implement a two-stage training recipe: first, full fine-tuning (FFT) on a small, high-quality in-domain corpus (e.g., 28 hours) for maximum accuracy; second, optional parameter-efficient fine-tuning (PEFT) with high-rank rsLoRA (r=128, α=256) to scale to larger models under memory constraints.
-
The improved system can achieve state-of-the-art WER (23.44%) on Burmese medical speech, outperforming general-domain models by >9% absolute, while offering a memory-efficient alternative for large models (e.g., Whisper-Large-v2) with only a 72% WER reduction from zero-shot baseline.
- Controlled Multi-Level Data Augmentation for Robustness
-
Integrate a selective augmentation strategy: apply waveform-level (time shift, pitch shift, Gaussian noise) and spectrogram-level (time/frequency masking) to 50% of training data, with simulated room acoustics (L-shaped, random-distance) via Pyroomacoustics.
-
The improved system can maintain high accuracy on clean speech (minor penalty) while reducing WER by up to 15–20% in noisy (SNR < 10 dB) and reverberant conditions, making it deployable in real clinical environments with variable acoustics.
- Phonological Error-Aware Post-Processing
-
Use the error analysis (voiced/voiceless plosive confusion, weak-prefix deletion, function-word omission) to build a pronunciation lexicon and a grammatical particle restoration model.
-
The improved system can automatically correct common Burmese medical misrecognitions (e.g., /bé/ → /pé/) and re-insert dropped nominalizing prefixes (/P@/) and polite markers (/bà/), reducing residual errors by an estimated 30–40% without retraining.
- Acoustic Robustness Scoring for Model Selection
-
Develop a meta-learner that predicts optimal augmentation settings per model size and target environment (clean vs. noisy vs. reverberant) based on SNR and room impulse response characteristics.
-
The improved system can dynamically switch between augmented and non-augmented checkpoints, achieving best-of-both-worlds performance: 23.44% WER on clean and near-augmented-level robustness in noisy settings.
- Latency-Aware Deployment Scheduler
-
Leverage the observed RTF differences (PEFT > FFT) to build a runtime scheduler that routes short utterances to FFT models and long, complex utterances to PEFT models, balancing accuracy and real-time constraints.
-
The improved system can maintain low-latency responses (< 1.0 RTF) for interactive clinical use while preserving high accuracy on difficult, longer medical dictations.
- Cross-Lingual Transfer for Tonal Medical Loanwords
-
Use the Burmese tonal pitch contour analysis to create a feature extractor that aligns with English medical loanwords (e.g.,
diabetes,
hypertension
) in low-resource languages. -
The improved system can generalize to other tonal languages (e.g., Thai, Vietnamese) with minimal fine-tuning, improving ASR for medical terms by 10–15% WER over vanilla Whisper.
- Simulated-to-Real Domain Adaptation
-
Incorporate a domain-adversarial training step that uses synthetic room impulse responses as a source domain and real clinical noise profiles (e.g., babble, device fan) as a target domain, using a gradient reversal layer.
-
The improved system can bridge the gap between simulated and real-world acoustics, reducing WER by an additional 5–8% in actual hospital settings compared to models trained only on synthetic augmentation.
Abstract
Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by native speakers. We fine-tune Whisper models using full fine-tuning (FFT) and parameter-efficient fine-tuning (PEFT) with LoRA. To evaluate robustness, we apply waveform- and spectrogram-level data augmentation under controlled noise and simulated room acoustics. While augmentation reduces performance on clean speech, it significantly improves robustness in noisy and reverberant environments across FFT and PEFT settings. Our best-performing system, fully fine-tuned myMediWhisper-Medium without augmentation, achieves a state-of-the-art Word Error Rate (WER) of 23.44%, outperforming much larger general-domain fine-tuned models. Dataset and other resources can be found at the Huggingface repository: https://huggingface.co/datasets/LULab/mediTalk-mm-rdy.
Sources
- FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering