V= a kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
cs.CL, eess.AS
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: Paper is accepted in IEEE SLT 2026
Code: https://github.com/linto-ai/whisper-timestamped
License: http://creativecommons.org/licenses/by/4.0/
The gist: Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings.
Terminology
Abstract
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.
Sources
- A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering
- ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering