Evaluating Large Language Models on Urdu Idioms

arXiv:2510.17460 · cs.CL · Submitted 2025-10-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evaluating Large Language Models on Urdu Idioms".

Jane: The paper was written by Muhammad Farmal Khan and Mousumi Akter from Technical University of Dortmund, Germany and @tu-dortmund.de email domain for affiliation purposes (not listed as a formal organization).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Methodology and Scope: Jane: So, "Evaluating Large Language Models on Urdu Idioms" sets the stage by explaining exactly what they did to conduct their tests, which is where the methodology comes in.

Tom: They didn't just test one single model, Jane; they evaluated a wide range of open-source Large Language Models to see how different architectures performed.

Lu: What's impressive is that they paired the LLM tests with traditional Neural Machine Translation or NMT systems to create a true benchmark for cutting edge AI capabilities.

Meng: The methodology here is robust; running experiments on two specific datasets, the Native Urdu dataset and the Roman Urdu dataset, allows us to compare formal versus informal data structures.

Lalam: These datasets are key because they allow us to observe the difference between highly structured linguistic forms and how people actually communicate in everyday settings.

Tom: It’s clear that by testing both LLMs and NMT on these specific sets, they are building a comprehensive picture of the current state of AI capabilities in Urdu.

Jane: They used several automatic metrics like BERTScore, COMET, and XCOMET to measure the quality of the translations, which is a sophisticated way to assess success beyond simple word overlap.

Lu: Those metrics go beyond simple word overlap; they look at deeper semantic equivalence, which is exactly what's needed for idiomatic translation where the meaning changes completely.

Meng: And when you look at Table two in the paper, it’s clear that even with the data collection effort, we have a solid baseline of one thousand one hundred and four hundred sixty idioms to compare against.

Lalam: The commitment to comparing these two writing systems—native and Roman—is what makes this study so valuable for understanding how language is used in practice across all regions.

Tom: This careful setup ensures that the results aren't just a snapshot but a highly detailed comparison of how we approach translation tasks.

Key Findings and Improvements: Jane: Now that we know the methodology, let's talk about what "Evaluating Large Language Models on Urdu Idioms" revealed in terms of performance and improvements.

Tom: The authors found a significant advantage for Large Language Models over traditional NMT systems across almost all metrics, which is a huge win for LLMs.

Lu: This suggests that the reasoning capabilities of modern LLMs are much better at understanding linguistic nuance than previous AI models were, which is a major leap forward in terms of cognitive ability.

Meng: And the practical implication here is that if we need to translate culturally sensitive or idiomatic content, we're leaning towards these larger, instruction-tuned LLMs for reliable output.

Lalam: I found the results on prompt engineering particularly insightful; guiding the model with specific instructions makes a noticeable difference in how well it captures that cultural meaning.

Tom: It’s not just about the model type either way, Jane. The paper clearly shows that using tailored prompts—like Cultural or Paraphrase prompts—consistently outperforms a simple Literal prompt.

Jane: That means we're learning how to *talk* to the AI, Tom, to get better results; we have to tell it exactly what kind of translation we want instead of just hoping it guesses correctly.

Lu: And while the performance differences between different types of LLMs are relatively minor, the impact of guiding them with prompts is not insignificant at all.

Meng: The data also showed a consistent trend where Native Urdu inputs produced significantly more accurate translations than those written in Roman Urdu, which is a practical finding.

Lalam: That's a vital finding for cultural preservation; if the input representation is inconsistent, the output accuracy suffers, highlighting how important standardized orthography is for accuracy.

Tom: So, we're seeing that both prompt engineering and also having a better source format can really boost the quality of idiomatic translation.

Jane: It’s a double win for making machine translation more effective and culturally sensitive by using LLMs to capture cultural context.

Implications and Future Impact: Tom: We've covered the methodology, the results, and how "Evaluating Large Language Models on Urdu Idioms" is changing our understanding of AI translation capabilities.

Jane: It’s a powerful demonstration that LLMs are moving beyond just being a dictionary replacement; they are becoming cultural interpreters who understand context.

Lu: I can already picture the possibilities for future work in other low-resource languages, applying this same techniques to build robust systems for Persian or Hindi based on these findings.

Meng: From an engineering viewpoint, this confirms that the investment in better prompt design and targeted evaluation is paying off in terms of higher quality output.

Lalam: The vision I have is that we are moving toward AI that understands the *context* of a world, not just its words, and this paper makes that step much clearer for global communication.

Tom: It's a clear benchmark for the future—it gives other researchers something concrete to measure against when they tackle their own translation problems.

Jane: We've seen how the Native Urdu input outperforms Roman Urdu, which is a crucial detail about consistency in language use that we need to keep in mind.

Lu: And I think this also paves the way for further research into how these models handle even more complex idiomatic structures moving forward with new data sets.

Meng: By showing that LLMs excel at semantic alignment, it suggests we can build practical applications that are much more reliable than previous systems we've used.

Lalam: It’s exciting to see AI capable of understanding the nuances of cultural idioms, making this a big step for better global communication and cultural exchange.

Wrap-up and Farewell: Tom: As we wrap up our discussion on "Evaluating Large Language Models on Urdu Idioms," I think we can all agree that this is a massive step forward for language technology.

Jane: It’s wonderful to see the clear evidence of how prompt engineering helps guide powerful AI toward cultural accuracy, which is such an important finding.

Lu: This sets a high bar, Tom; we have so much more to learn about the limits and potential of these models as research continues.

Meng: The practical impact is that we're building more trustworthy and usable AI systems for real-world applications that need reliable translation quality.

Lalam: I hope this work helps bridge the cultural gap in translation across different languages globally, making communication smoother for everyone involved.

Tom: It's clear that we're leaving a lot of ground covered here, but also opening up so many doors for future research to explore.

Jane: We'll be sure to look at those next steps and dive deep into the next set of papers on the horizon with this same level of scrutiny.

Lu: The potential for is limitless; it’s just the beginning of a new era in how we process and understand human language.

Meng: For us, it means that building systems that respect cultural integrity is becoming paramount for real-world AI development.

Lalam: It's a moment of deep cultural progress, and I am truly thrilled to have shared this discussion with you all today.

Technical University of Dortmund, Germany · @tu-dortmund.de email domain for affiliation purposes (not listed as a formal organization)

cs.CL

Submitted: 2025-10-20

Updated: 2026-09-03

Comments: Accepted to Findings of EMNLP 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: The paper evaluates multiple open-source LLMs and NMT models for translating idioms from both Native and Roman Urdu.

Key concepts

Prompt Engineering
Guiding a model with specific instructions is crucial for better results. The paper found that using tailored prompts, such as 'Cultural' or 'Paraphrase' prompts, consistently produces more accurate translations than simply using a literal prompt.
Native vs. Roman Urdu
The study used two datasets to compare how language is represented. Native Urdu inputs produced significantly more accurate translations than those written in Roman Urdu, highlighting the importance of standardized orthography for accuracy.

Terminology

Summary

The paper evaluates multiple open-source LLMs and NMT models for translating idioms from both Native and Roman Urdu. This research establishes a crucial benchmark for low-resource idiomatic translation in Urdu by testing model capabilities across various semantic metrics, thereby demonstrating the effectiveness of advanced techniques like prompt engineering combined with powerful instruction-tuned LLMs.

Model Performance Benchmarking

The study rigorously tested model performance using four primary automatic evaluation metrics:

  • (a) BERTScore comparison

  • (b) BLEU comparison

  • (c) COMET comparison

  • (d) XCOMET comparison

The results show that the GPT-OSS-20B model consistently outperforms other models across semantic metrics (BERTScore, COMET, XCOMET). This superior performance indicates strong fluency and cultural understanding in idiomatic translation. The evaluation confirms that instruction-tuned large models are highly effective for this specialized task.

Impact of Input Format and Prompt Design

Translation quality is significantly influenced by both the input script and the prompt structure provided to the LLMs. Regarding prompting, the paper notes that:

  • Cultural and Paraphrase prompts generally improve semantic alignment over Literal prompts.

  • Conversely, smaller models exhibit limited sensitivity to prompt variations, suggesting that larger models benefit more from sophisticated prompting techniques.

Furthermore, when comparing the two writing systems used in the study, Native Urdu inputs consistently yield higher translation quality compared to Roman Urdu, which highlights the importance of orthographic representation for accurate machine translation.

Study Limitations

While providing a valuable benchmark, the authors acknowledge several limitations that restrict the generalizability of their findings. These include:

  1. Dataset Size: The datasets are described as relatively small, comprising only 1,100 Native Urdu idioms and 460 Roman Urdu idioms, which may limit findings to broader contexts.

  2. Orthographic Variability: The lack of standardized spelling and orthography in Roman Urdu introduces considerable variability, potentially skewing evaluation outcomes.

  3. Evaluation Scope: The assessment relies exclusively on automatic metrics (BERTScore, BLEU, COMET, and XCOMET), which may fail to fully capture the deep idiomatic meaning or cultural nuance inherent in the language.

  4. Model Exclusion: The study assessed only open-source models and explicitly excludes proprietary or larger-scale models that might yield stronger results but are not publicly available.

Improvements for AI systems

Based on a meticulous review of this research—particularly its focus on low-resource, culturally nuanced, idiomatic translation across divergent orthographies (Native vs. Roman Urdu)—several critical advancements are necessary to move beyond current benchmarks and build truly robust AI systems.

My proposed improvements focus on three core pillars: Orthography Normalization & Contextual Embedding, Deep Semantic Modeling, and Adaptive Prompting Architectures.


The most significant vulnerability exposed is the variability and lack of standardization in Roman Urdu. Current systems treat variability as noise; they must treat it as a solvable linguistic problem.

  • Improvement: Develop and integrate a dedicated, pre-processing module—the Cross-Lingual Orthographic Normalization (CLON) layer. This layer must map all known Roman Urdu variations of the same phoneme/word back to a canonical, standardized Unicode representation before feeding the text into the core NMT or LLM pipeline. This requires building a comprehensive, language-specific grapheme-to-phoneme (G2P) dictionary trained on massive parallel corpora that includes transliteration guides and common user errors.

  • What the Improved AI System Can Do: It will achieve Robust Input Equivalence. The system will process inputs like 'kya haal hai,' 'kia hal hai,' and 'kyah haal hai' not as three separate inputs, but as a single, canonical embedding representing How are you? This drastically reduces the variance-induced error rate (E var) and allows the core translation model to focus purely on semantic meaning rather than orthographic ambiguity.

The current reliance on standard metrics (BERTScore, COMET) is insufficient because they often fail to quantify deep cultural resonance or idiomatic divergence—the cultural nuance mentioned in the conclusion.

  • Improvement: Implement a Culturally Grounded Semantic Embedding Space (CGSE) that augments the standard embedding layer. This requires fine-tuning an existing LLM (e.g., Mistral or Gemma) not just on parallel text, but on curated datasets of cultural context explanations and semantic divergence examples. The CGSE must be trained to identify the literal meaning vector versus the intended idiomatic meaning vector.

  • What the Improved AI System Can Do: It will achieve Intentional Semantic Transfer. When translating an idiom like 'Aankhon ka taara' (literally 'star of the eyes'), instead of simply mapping it to a high-similarity English phrase, the system will map it to a vector space representing its function within Urdu culture (e.g., cherished loved one). This allows for contextually appropriate paraphrasing that respects cultural weight, which is critical for generating high-quality output in low-resource settings.

The paper notes that prompt design is crucial, but this suggests a manual, ad-hoc process. This needs to be codified into an active architectural component.

  • Improvement: Develop the Adaptive Prompting and Retrieval Augmentation Framework (APRAF). This framework sits around the core LLM call. Before translation, it dynamically queries a highly specialized, indexed knowledge base (KB Idioms) containing: 1) Canonical Idiom Pairings (Urdu English); 2) Known Prompt Strategies (Cultural, Literal, Paraphrase); and 3) Orthography Maps (from CLON). Based on the input characteristics (e.g., high Romanization variance detected by CLON), APRAF automatically selects and constructs the optimal prompt template and retrieves the most relevant few-shot examples from KB Idioms to prime the LLM.

  • What the Improved AI System Can Do: It will achieve Self-Optimizing Translation Pipeline. The system moves beyond simply receiving a prompt improvement; it determines and applies the necessary prompt improvement autonomously. If the input is ambiguous, it automatically asks for clarification or defaults to a safer, highly constrained paraphrasing mode, drastically reducing the risk of generating semantically inconsistent translations.


Component Problem Addressed Core Mechanism Achievable Capability

:---:---:---:---

CLON Module (Input) Roman Urdu variability (E var) and low standardization. Grapheme-to-Phoneme mapping to a canonical Unicode form. Robust Input Equivalence: Translating based on sound/phoneme, not spelling variation.

CGSE Layer (Model Core) Failure to capture deep cultural meaning or idiomatic function. Fine-tuning LLMs on semantic function vectors (vs). Intentional Semantic Transfer: Generating translations that resonate culturally, not just linguistically.

APRAF Framework (Control) Reliance on manual prompt engineering; non-adaptivity. Dynamic querying of a knowledge base to construct optimal prompts and few-shot examples at runtime. Self-Optimizing Pipeline: Automatically selecting the best translation strategy based on input complexity and ambiguity.

Abstract

Idioms remain a persistent challenge in natural language processing due to their figurative and culturally grounded meanings, which distinguish them from literal expressions. Although recent advances in large language models (LLMs) have improved idiom handling across several languages, limited attention has been given to low resource languages such as Urdu. In this work, we present a comprehensive benchmark for Urdu to English idiomatic translation, consisting of a manually verified dataset of 4,000 aligned idiom sentence pairs in both Perso Arabic (native Urdu script) and Romanized Urdu. We evaluate multiple tasks, including translation, paraphrasing, idiom span detection, and back-translation, using diverse prompting strategies such as cultural prompting, idiomatic prompting, and few-shot learning. Our findings show that frontier LLMs consistently outperform traditional neural machine translation systems across all evaluation settings, particularly in preserving figurative and metaphorical meaning. While models demonstrate relatively stable performance on native Urdu script, the absence of standardized orthography in Romanized Urdu introduces substantial challenges for consistency and idiom span detection. This work establishes a high quality benchmark for cross script idiomatic evaluation in Urdu and underscores the importance of prompt engineering in preserving figurative language meaning across languages.

Sources

Related papers