Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages

arXiv:2608.29239 · cs.CL, cs.SD, eess.AS · Submitted 2026-08-29 · Read on arXiv

cs.CL, cs.SD, eess.AS

Submitted: 2026-08-29

Updated: 2026-08-29

Comments: Accepted to EMNLP 2026 (Main Conference)

License: http://creativecommons.org/licenses/by/4.0/

The gist: Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation.

Terminology

Abstract

Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation. To address this gap, we propose SAMA-ASR, a lightweight adapter mechanism that augments the decoder with semantic anchors from auxiliary translations and an acoustic anchor from speech; in principle, the mechanism can be applied to similar encoder--decoder multitask speech models. Through cross-modal adaptation, SAMA-ASR conditions decoder states on translation-derived semantic embeddings and a speech embedding, combining utterance-level meaning with speech-grounded evidence before token prediction. At evaluation time, these semantic anchors can be generated automatically by an upstream speech-to-text translator rather than supplied as oracle translations. Experiments on two 30-hour datasets covering the low-resource Sinitic varieties Taiwanese Hokkien and Hakka show that SAMA-ASR improves over acoustic, prior prompt-based, and semantic-only translation-guided baselines and remains effective in practical automatic semantic-anchor settings; translator-capacity analyses show that useful semantic anchors can be produced by a compact ST model.

Related papers