Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services

arXiv:2609.09889 · cs.CL · Submitted 2026-09-09 · Read on arXiv

cs.CL

Submitted: 2026-09-09

Updated: 2026-09-09

Comments: Published in Engineering Applications of Artificial Intelligence

Journal ref: Engineering Applications of Artificial Intelligence 183 (2026) 116038

DOI: 10.1016/j.engappai.2026.116038

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription.

Terminology

Abstract

Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments such as call center conversations. When privacy restrictions preclude audio access, error correction must rely on text-based post-editing. Existing text-only approaches face significant challenges in low-resource languages, mainly due to a critical scarcity of annotated corpora and tailored correction methodologies. For Korean, this resource gap is particularly pronounced, as existing resources are predominantly designed for ASR training rather than text-based error correction. To address this, we introduce DasanCallDial, the first large-scale Korean benchmark dataset specifically curated for dialogue-level ASR error correction. Derived from genuine call center interactions, it comprises 1,974 dialogues with 115,460 utterances. Leveraging this resource, we propose Detector-Gated Contextual Span Correction (DCSC), a text-only post-editing framework for error-sparse Korean speech recognition transcripts. DCSC combines an encoder-based detector that first performs token-level error detection, followed by a language model-based corrector trained to rectify fine-grained span-level errors. Additionally, we employ dialogue-level context augmentation to enable the model to leverage discourse history for disambiguation. By employing multi-level granularity, our method achieves state-of-the-art performance, effectively overcoming the limitations of general LLMs in low-resource settings.

Related papers