MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
cs.CL
Submitted: 2025-08-19
Updated: 2026-03-18
Comments: Accepted to LREC 2026
Journal ref: https://aclanthology.org/2026.lrec-1.334/
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: In this paper, we introduce MATA, a novel evaluation dataset to assess the ability of Large Language Models (LLMs) in Telugu language, comprising 729 carefully curated multiple-choice and open-ended
Terminology
Abstract
In this paper, we introduce MATA, a novel evaluation dataset to assess the ability of Large Language Models (LLMs) in Telugu language, comprising 729 carefully curated multiple-choice and open-ended questions that span diverse linguistic dimensions. We evaluate 11 open-weight and closed-source LLMs on our dataset and present a fine-grained analysis of their performance. Further, we empirically show how LLMs rely on superficial heuristics such as answer position and distractor patterns for multiple-choice questions. Finally, we also compare LLM-as-a-judge evaluation with human evaluation for open-ended questions assess its reliability in a low-resource language. We argue that such fine-grained evaluation is essential for understanding model limitations and can inform the development of more linguistically capable LLMs, while also serving as a foundation for future research in Telugu NLP. Our dataset is available at: https://huggingface.co/datasets/TeluguLLMResearch/MATA
Sources
- NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts
- Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
- Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
- Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
- HKCanto-Eval: A Benchmark for Evaluating Cantonese Language Understanding and Cultural Comprehension in LLMs
- Evaluating Polish linguistic and cultural competency in large language models
- IndicSQuAD: A Comprehensive Multilingual Question Answering Dataset for Indic Languages
- TLUE: A Tibetan Language Understanding Evaluation Benchmark
- BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
- Evalita-LLM: Benchmarking Large Language Models on Italian
- Spanish and LLM Benchmarks: is MMLU Lost in Translation?
- Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
- INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
- None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks
- KoBALT: Korean Benchmark For Advanced Linguistic Tasks
- From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation
- SEA-HELM: Southeast Asian Holistic Evaluation of Language Models
- IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
- MILU: A Multi-task Indic Language Understanding Benchmark
- F\`ux\`i: A Benchmark for Evaluating Language Models on Ancient Chinese Text Understanding and Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering