Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR

arXiv:2605.29637 · cs.CL · Submitted 2026-05-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR".

Jane: Large language models often struggle with cross-lingual knowledge consistency when querying in low-resource languages, especially Indian languages and their code-mixed variants.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, what is the actual finding they reported in terms of performance? I’m interested in knowing exactly how big this gap between native and English inputs is supposed to be, and what happens when they introduce code-mixed inputs.

Jane: Well, the summary highlights a pronounced consistency gap between native and English inputs on the INDIKLAR benchmark, with some discrepancies reaching up to zero point five zero in accuracy. But what’s surprising is that when they use code-mixed variants, that gap substantially closes without needing any special changes to the model itself.

Lu: That suggests the model has some kind of latent ability to handle the mixed input structure internally, which connects back to ideas about bilingual cognition where speakers often process things in their dominant language. It implies that the structure of the code-mixed text itself is helping it navigate the knowledge space.

Meng: If it closes without model intervention, that’s fantastic for deployment because we don't need to spend massive resources retraining or fine-tuning just to handle this linguistic variation. It suggests an implicit mechanism is working well.

Lalam: It means that for Indian languages, the way we naturally speak and write together might actually be a pathway for the AI to access its knowledge more efficiently than if it only saw pure native text. This has implications for how we design future conversational agents.

The paper's summary: Tom: Okay, moving on, what about the strategies they tested to actually fix this gap? They tried a few prompting methods, and I want to hear what they found was actually effective for improving consistency.

Jane: They explored several prompting strategies, including a two-stage approach and a one-stage joint translation method. But the most effective one they identified was Translate-in-Thought, or TinT, which is a single step where the model does the conversion internally and only gives the final answer.

Lu: TinT is powerful because it seems to be activating those internal mechanisms we were talking about—like better neuron utilization—without needing massive retraining. It shows that activating the right internal behavior through prompting alone can substantially improve how existing multilingual knowledge is used.

Meng: I’m looking at the practical side here; if TinT offers gains, we need to see how much time it adds to the inference process compared to just answering directly. If it’s a single step, that’s manageable for real-time applications.

Lalam: It really shifts the focus from changing the model weights to optimizing how we ask the model questions—which feels much more flexible when dealing with eighteen different languages. This gives us a tool we can deploy right away.

The paper's improvements: Tom: So, to wrap up this paper on "Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR," what is the big picture message they are sending about these low-resource languages?

Jane: The main conclusion is that code-mixed inputs significantly bridge the gap between native and English performance, and that prompting strategies like TinT are quite effective at recovering accuracy on their own. They found a consistent flip point where predictions start getting correct somewhere between the native language setting and the code-mixed one.

Lu: That flip point is crucial because it tells us *where* the boundary is, whether it's due to how we input things or how the model processes them internally. It helps us pinpoint exactly what kind of linguistic structure unlocks better reasoning in these contexts.

Meng: For practical implementation, this means if we’re building tools for Indian languages, focusing on code-mixed examples might give us a quicker path to reliable results than trying to build perfect monolingual models from scratch right away.

Lalam: I think what this study really shows is that we don't always need huge amounts of data or massive retraining to get good cross-lingual recall, especially when we use smart prompting techniques to guide the model’s internal reasoning.

Tom: That sounds like a strong summary. It seems like the future here is less about brute force and more about clever instruction design for these complex language scenarios. We’ve got some solid insights into how we can start thinking differently about cross-lingual consistency across diverse languages.

Conclusion: Tom: So, to wrap up, this paper on "Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR" really shows that we can actually see how language structure helps or hinders an AI's ability to recall facts across different languages.

Jane: Exactly. The main thing is that code mixing doesn't just introduce noise; it surprisingly closes the gap between native and English inputs without us needing to retrain anything, which is a really neat observation for how these models learn.

Lu: I think what’s most exciting about this work is identifying that flip point; it gives us a clear boundary on when the model's internal processing starts favoring one input style over the other, whether it’s native script or code-mixed text. That suggests a deeper understanding of how knowledge is organized in these multilingual systems that we can use to build better architectures.

Meng: From an engineering standpoint, this means we might not need to build entirely separate pipelines for every single low-resource Indian language if the code-mixed approach proves robust, which could simplify our deployment strategy immensely.

Lalam: I feel like the biggest impact here is on how we think about cultural knowledge recall; if we can get an AI to understand concepts through these natural, mixed forms, it opens up ways for AI to interact with and reflect our own diverse linguistic heritage in a much more meaningful way.

Tom: That's a powerful vision, Lalam. It really puts the focus on leveraging existing linguistic patterns instead of just throwing more data at it, which is what this research suggests.

Jane: And the next big thing we can look toward is how to integrate those findings into real-time applications so that users get consistent, reliable answers across all these language variants without any extra latency.

Lu: I’m curious if there are future papers that tackle the scaling of this effect; we need to see how this consistency holds up when we move from smaller models to those with much larger parameter counts <ref:two thousand six hundred five point two nine six three seven#pg2.

Meng: Scaling is always the hurdle for me; if the TinT prompting strategy works well across different model sizes, that’s a huge win for resource allocation in building scalable multilingual systems.

Lalam: I hope we see more research focusing on how these cross-lingual connections can help build AI that genuinely understands the nuances of human culture through language, not just translating words <ref:two thousand six hundred five point two nine six three seven#pg1.

Tom: Absolutely. We’ve got some serious material here with this IndicKLAR study, and it makes me really look forward to what comes next in this field.

Debajyoti Mazumder, Divyansh Pathak, Prashant Kodali, Aditya Joshi, Akshay Agarwal, Jasabanta Patro

Indian Institute of Science Education and Research, Bhopal, India · Microsoft Corporation

cs.CL

Submitted: 2026-05-28

Updated: 2026-09-29

Importance score: 82/100

The gist: Large language models often struggle with cross-lingual knowledge consistency when querying in low-resource languages, especially Indian languages and their code-mixed variants.

Key concepts

INDIKLAR
A new benchmark designed specifically for evaluating AI model performance across 18 Indian languages and their code-mixed variants paired with English. It allows researchers to compare how models perform on native text, mixed text, and pure English inputs simultaneously.
Code-Mixed Inputs
Text that mixes words from an Indian language with words from another language, typically English. These are created by inserting English content words into the grammatical structure of a specific Indian language to simulate real-world usage.
TinT Prompting
A single-step prompting strategy where the model internally translates the input before generating its final answer. This method was found to be highly effective, significantly improving cross-lingual knowledge recall compared to standard translation or multi-stage prompting methods.

Terminology

Summary

Large language models often struggle with cross-lingual knowledge consistency when querying in low-resource languages, especially Indian languages and their code-mixed variants. This work introduces INDIKLAR, a new benchmark that evaluates model performance across native Indian languages, code-mixed forms, and English inputs to examine this gap. The key finding is that while there is a significant accuracy gap between native and English inputs (up to 0.50), code-mixed inputs substantially close this gap without any model-level intervention.

Indiklar Benchmark Introduction

The paper introduces INDIKLAR, an Indic extension of KLAR-CLC covering 18 of the 22 scheduled Indian languages and pairing them with code-mixed variants for 11 widely used language pairs. This benchmark is unique because it provides a three-way alignment, allowing evaluation across English, native Indian languages, and naturally motivated code-mixed inputs. The code-mixed variants are constructed by interleaving English content words into the grammatical structure of each Indian language. The dataset statistics show that INDIKLAR provides threeway parallel variants (monolingual native Indian, code-mixed, and English) for 11 languages, with 2,619 instances spanning 20 fact categories per language setting.

Performance Gap Analysis

The evaluation compares model robustness across three input settings: (i) English (EN), (ii) Native (NA), and (iii) Code-Mixed (CM). The results reveal a pronounced consistency gap between native and English inputs—one that, surprisingly, code-mixed variants largely close without any model-level intervention. Specifically, the Native tier exhibits severe degradation compared to English, while CM scores consistently approach English-level performance with a mean gap remaining within ∼0.05. This indicates that code-mixing effectively bypasses the low-resource penalty at inference time.

Prompting Strategies and TinT

The study evaluates several prompting strategies designed to bridge the crosslingual gap, including a two-stage translate-then-answer setup, a one-stage joint translation-and-answer prompt, and Translate-in-Thought (TinT)—a singlestep strategy in which the model converts the input internally and emits only the final answer. TinT is shown to be highly effective, showing statistically significant gains over the baseline across languages and model families. The paper demonstrates that TinT achieves competitive gains for Cross-Lingual Consistency (CLC), with a gain of +0.1492, outperforming other strategies in this metric.

Flip Point Identification

A central finding is the identification of a consistent flip point—the boundary between incorrect and correct prediction—that lies between the native and code-mixed settings. This boundary holds whether it is induced by the input surface form or by the model’s internal conversion process. The analysis shows that Gold code-mixed inputs (explicit evaluation) bring performance close to English-level accuracy, while TinT prompting (implicit evaluation) independently recovers a significant portion of the same gap.

Key Methodological Insights

The research employs several rigorous methods to isolate the effects of language variation and prompting. The methodology includes:

  1. Evaluating models from multiple families: Llama 3.1, Gemma 3, and Qwen 2.5, spanning roughly 1B–14B parameters.

  2. Using two primary metrics: Accuracy (defined via prefix-matching) and Cross-Lingual Consistency (CLC), which is computed using the overlap ratio between correctly predicted sample indices across language pairs.

  3. Ablating the effect of code-mixing versus romanization, finding that code-mixed variants substantially outperform both native-script (low-res.) and romanized baselines, suggesting gains stem from meaningful bilingual lexical cues rather than script conversion alone.

  4. Analyzing layer-wise rank using a Logit Lens analysis to track the gold answer token's rank across model layers, revealing that the gap between TinT and baseline is concentrated in the deeper layers, consistent with late-layer cross-lingual reorganization.

Conclusion and Generalization

The study concludes that TinT prompting demonstrates broad effectiveness beyond INDIKLAR to a contextually mediated knowledge recall benchmark, suggesting that implicit translation prompting is broadly effective for crosslingual factual knowledge recall. The findings suggest the flip point lies between native and code-mixed settings in the performance trajectory of low-resource → codemixed → English. Furthermore, TinT’s effectiveness scales with model size, as larger models possess stronger latent translation ability, while Gemma-3-1B bucks the small-model trend by improving across all languages. The study also notes limitations, such as synthetic generation of variants and evaluation on models up to 14B parameters.

Improvements for AI systems

Here are specific, actionable improvements for AI systems derived from the insights in this paper, focusing on enhancing cross-lingual knowledge consistency, particularly for low-resource and code-mixed Indian languages:


)1. Implement a Code-Mixed Consistency Layer (Inspired by INDIKLAR/TinT):

Instead of relying solely on direct translation or native script input for low-resource languages, integrate a latent code-mixing mechanism into the inference pipeline.

The system should be prompted to internally convert the query into a structured English/Roman code-mixed form before answering (TinT-CM).

This allows models to leverage their superior English knowledge base while maintaining grammatical coherence in the target Indian language script, effectively bypassing low-resource penalties.

)2. Develop Adaptive Prompting Strategies Based on Language Tiering:

Design a dynamic routing mechanism for inference based on the input language (Native vs. Code-Mixed vs. English).

If the input is code-mixed, prioritize the 1Step-CM+Ans strategy to maximize lexical grounding benefits without incurring high latency.

If performance lags significantly below English baseline (as seen in Figure 5), automatically switch to a multi-step translation/answer setup (like 2Step-EN) for maximum accuracy recovery, while monitoring for inference time constraints.

)3. Optimize Model Architecture and Fine-Tuning via Latent Translation Focus:

Investigate fine-tuning objectives that explicitly encourage internal translation capabilities rather than just surface pattern matching.

Fine-tune models using auxiliary loss functions that reward the correct internal representation (as measured by the layer-wise rank analysis in §6.1) rather than just maximizing token probability on the final output. This aims to bridge the gap between base and instruction-tuned models for cross-lingual reasoning.

)4. Enhance Low-Resource Language Coverage via Synthetic Code-Mixing:

Address the limitation that only 11 languages have human verification by using GPT APIs to generate high volumes of syntactically correct, naturalistic code-mixed variants for the remaining 7 languages (Dogri, Kannada, etc.).

This synthetic data should be used to create a robust Code-Mixed Consistency Layer for these under-represented languages, enabling their inclusion in cross-lingual benchmarks like INDIKLAR.

)5. Implement Contextually Mediated Recall Benchmarks:

Move beyond simple entity queries (like KLAR-CLC) and adapt evaluation frameworks to test knowledge recall within naturalistic, referential contexts (as explored in §5.7).

The AI system should be tested on contextually mediated benchmarks where the knowledge required is inferred from a larger surrounding text, simulating real-world usage rather than isolated factual lookups.

)6. Introduce Latency-Aware Prompting for Production Deployment:

Prioritize inference efficiency by defaulting to TinT variants unless a specific, high-accuracy requirement dictates otherwise.

The system should use the TinT variant (TinT-CM or TinT-EN) as the default production setting due to its minimal latency overhead compared to multi-step strategies, ensuring real-time performance for multilingual applications.

Sources

Related papers