Automatic Speech Recognition for Multilingual Oral History Research

arXiv:2609.04232 · eess.AS, cs.CL · Submitted 2026-07-21 · Read on arXiv

eess.AS, cs.CL

Submitted: 2026-07-21

Updated: 2026-07-21

Comments: This preprint reflects an updated version of the manuscript prepared for Interspeech 2026, incorporating revisions based on reviewer feedback

License: http://creativecommons.org/licenses/by/4.0/

The gist: This paper offers a unique perspective on how speech technologies are being adopted by community-led heritage language preservation and revitalisation initiatives.

Terminology

Abstract

This paper offers a unique perspective on how speech technologies are being adopted by community-led heritage language preservation and revitalisation initiatives. As a community-led language maintenance strategy, oral histories play a crucial role in Cantonese language revitalisation in New Zealand. The development of Automatic Speech Recognition (ASR) toolkits, such as Whisper, have expedited what has often been a resource and time-intensive process of transcribing oral history collections. However, there is limited research into the effectiveness of ASR toolkits when applied to code-switched language contexts. Based on Word Error Rate (WER), the best performing Whisper model configuration achieved a WER of 12.10 at the expense of accurately transcribing unsupported non-English segments. However, Whisper remains a useful tool by providing a first-pass transcription using only 1% of the estimated time otherwise needed for manual transcription.

Related papers