HearInContext: A Benchmark for Implicit Context in Speech Recognition
cs.CL, cs.SD
Submitted: 2026-09-16
Updated: 2026-09-21
Code: https://github.com/OPPO-Mente-Lab/HearInContext
License: http://creativecommons.org/licenses/by/4.0/
The gist: Contextual ASR can benefit from semantic cues or from target words explicitly provided in the context.
Terminology
Abstract
Contextual ASR can benefit from semantic cues or from target words explicitly provided in the context. We introduce HearInContext, a Mandarin--English benchmark that pairs shared synthetic speech with assistant replies supporting different interpretations. The benchmark comprises 3,764 semantic test cases built around homophones. Implicit contexts exclude candidate words; explicit contexts name the target. No-context and unrelated-context controls measure the benefit of relevant history and sensitivity to irrelevant history. Context-capable models benefit from implicit cues but achieve higher target recall with explicit hints. Fine-tuning Qwen3-ASR-1.7B improves implicit-context target recall by 11.0 and 11.5 percentage points in Mandarin and English, respectively, while absolute CER/WER changes on AISHELL-1 and LibriSpeech remain below 0.1 percentage points. Gains extend to explicit conditions excluded from fine-tuning and to Mandarin hotword recognition on real recordings.
Sources
- Prompting Large Language Models for Zero-Shot Domain Adaptation in Speech Recognition
- ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
- PROFASR-BENCH: A Benchmark for Context-Conditioned ASR in High-Stakes Professional Speech
- IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
- Qwen3-ASR Technical Report
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
- VIBEVOICE-ASR Technical Report
- Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering