Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Obscuring Data Contamination Through Translation".
Jane: The gist The translation into Arabic can suppress conventional contamination indicators while still allowing models to benefit from exposure, particularly those with stronger Arabic capabilities.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the show, everyone. Today we’re talking about a paper that tackles how we check if large language models are cheating on benchmarks, specifically data contamination. We’ve got "Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora."
Jane: It sounds like they're looking at a problem where English-only tests might be missing something important when we look at models that handle other languages.
Lu: Exactly. The paper argues that translation into a low-resource language, like Arabic, can actually act as a barrier to contamination signals because it suppresses the usual indicators we look for in English benchmarks <ref:2601.14994#pg1>.
Meng: So the core idea is that if you translate the benchmark, you might hide the memorization effect, but it doesn't remove it entirely <ref:2601.14994#pg2>.
Tom: Right. It’s about finding a way to detect contamination that works across different languages without just relying on English tests <ref:2601.14994#pg3>.
Jane: The paper investigates this by fine-tuning open-weight models on Arabic datasets and then testing them against the original English benchmarks <ref:2601.14994#pg1>.
Lu: They extend a method called Tested Slot Guessing by adding a choice reordering strategy and Min-K percent probability analysis to catch both behavioral and distributional signals <ref:2601.14994#pg3>.
Meng: And their finding is that translation into Arabic suppresses those conventional contamination indicators, which is interesting because it suggests the language shift itself hides the pattern <ref:2601.14994#pg3>.
Lalam: But they found that models still benefit from being exposed to contaminated data, especially those with stronger Arabic capabilities <ref:2601.14994#pg1>.
Tom: That leads us into the second part of the paper—how they propose a way to see this problem more clearly. The authors introduce Translation-Aware Contamination Detection, or TACD <ref:2601.14994#pg3>.
Jane: So TACD is a diagnostic procedure that compares signals across multiple versions of a benchmark instead of just English ones <ref:2601.14994#pg3>.
Lu: They construct several evaluation views for each instance, which involves the question, the answer choices, and the correct answer index <ref:2601.14994#pg3>.
Meng: They use two signals to measure this: Index Recall Rate and Cross-Lingual Consistency <ref:2601.14994#pg3>.
Tom: Index Recall Rate tells us if a model is relying on memorized index-level associations instead of actually reasoning about the content <ref:2601.14994#pg3>.
Jane: And Cross-Lingual Consistency measures how much that consistency increases as the exposure to contaminated data grows across translations <ref:2601.14994#pg3>.
Lu: The empirical findings across models like Llama-three point two-1B-Instruct, Gemma-three-1B-it, and Qwen3 show different behaviors depending on the poisoning level <ref:2601.14994#pg4>.
Tom: For Llama, they saw low Index Recall Rate across all poisoning levels but cross-lingual consistency increasing gradually with the poisoned exposure <ref:2601.14994#pg4>.
Jane: Qwen showed perfect cross-lingual consistency regardless of the poisoning level while keeping its Index Recall Rate near chance, which points to a model and prompt specific collapse <ref:2601.14994#pg4>.
Meng: Whereas Gemma showed a different pattern where the Index Recall Rate stayed near random baseline but cross-lingual consistency increased steadily with poisoning <ref:2601.14994#pg4>.
Tom: So what does this mean for us practically? It suggests that we shouldn't just treat translation as some kind of decontamination process for evaluation pipelines <ref:2601.14994#pg5>.
Jane: Instead, it means we need detection pipelines that explicitly compare signals across translated versions rather than just sticking to English tests <ref:2601.14994#pg5>.
Lu: It also suggests that when auditing for contamination, we should pair aggregate metrics with behavioral probes that are robust to changes in the surface form of the text <ref:2601.14994#pg5>.
Meng: If we're going multilingual, we have to move past just assuming English-only checks are enough to maintain fairness and reproducibility <ref:2601.14994#pg5>.
Tom: So, looking at the full picture of "Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora," the authors show that translation into Arabic can suppress standard contamination cues, but models still show signs of being trained on contaminated data if you look at it through a multilingual lens <ref:2601.14994#pg1>.
Jane: The paper’s title really captures the main point: how translating data obscures the contamination signals that we usually look for in English benchmarks <ref:2601.14994#pg3>.
Lu: It’s a deep dive into why existing methods fail when models are exposed to diverse, multilingual data settings <ref:2601.14994#pg2>.
Meng: It gives us a concrete way to think about how memorization can hide itself depending on the language you're looking at <ref:2601.14994#pg3>.
Tom: So, the implication is that as evaluation gets more multilingual, we need smarter ways to detect contamination than just looking at one language’s statistics <ref:2601.14994#pg5>.
Jane: That means future benchmarking has to be much more careful about how it handles language differences in its testing protocols <ref:2601.14994#pg5>.
Conclusion: Tom: So, basically, this paper shows that translating those benchmark tests into Arabic actually hides the usual signs of data contamination we look for in English benchmarks <ref:2601.14994#pg3>.
Jane: That’s a big idea, Tom. It means if we only check for cheating using English tests, we might miss problems happening in other languages entirely <ref:2601.14994#pg5>.
Lu: Exactly. The authors are setting up this whole framework called Translation-Aware Contamination Detection to see those signals across different versions of the benchmark instead of just one language <ref:2601.14994#pg3>.
Meng: So, if a model is fine with Arabic text, it might be exploiting contamination in a way that's invisible when you only look at English results <ref:2601.14994#pg3>.
Lalam: From my side, the LLaMA model showed that even with Arabic translation, we still see that reliance on memorized patterns if we look at the Index Recall Rate <ref:2601.14994#pg3>.
Tom: And they found a way to measure that through these two signals, Index Recall Rate and Cross-Lingual Consistency, which gives us a much richer picture <ref:2601.14994#pg3>.
Jane: It really changes how we think about evaluation pipelines. We can't just treat translation as a way to clean up data; it's like putting a filter on the signal <ref:2601.14994#pg5>.
Lu: The implication is that when we build these benchmarks, they have to be designed with multilingual contamination in mind from the start <ref:2601.14994#pg5>.
Meng: So, for someone building AI systems, it suggests you need to check your contamination audits across all the languages you're testing in <ref:2601.14994#pg5>.
Lalam: It’s about making sure our models are fair across different linguistic capabilities rather than just one specific language <ref:2601.14994#pg3>.
Tom: Yeah, so the title itself, "Obscuring Data Contamination Through Translation," really sums up the main challenge they're addressing here <ref:2601.14994#pg3>.
Jane: It’s a reminder that as AI gets more global, our way of checking for honesty in those models needs to get more global too <ref:2601.14994#pg5>.
Lu: And this work provides the tools to start building those better, multilingual detection methods from scratch <ref:2601.14994#pg3>.
Department of Electrical and Computer Engineering, Maroun Semaan Faculty of Engineering and Architecture, American University of Beirut
cs.CL, cs.AI
Submitted: 2026-01-21
Updated: 2026-10-08
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: The gist The translation into Arabic can suppress conventional contamination indicators while still allowing models to benefit from exposure, particularly those with stronger Arabic capabilities.
Key concepts
- Data Contamination
- This occurs when a model relies on memorized answers from the training data rather than learning true general knowledge. It undermines the validity of model evaluations because it makes models appear better than they actually are, as they simply 'cheat' by recalling specific examples.
- Translation-Aware Contamination Detection (TACD)
- This is a diagnostic procedure designed to find contamination in multilingual settings. It works by creating multiple evaluation views from the same data instance—one for the original English and others for its Arabic translations—and comparing the resulting signals to spot memorization effects that might be hidden by translation.
- Cross-Lingual Consistency (CLC)
- CLC measures how consistent a model's behavior is when tested across different language versions of the same prompt. When CLC increases with poisoned exposure, it suggests that the model develops a general understanding of the concept rather than just memorizing specific content.
- Index Recall Rate (IDR)
- IDR tracks whether a model relies on recalling specific index-level associations or reasoning based on content. High IDR suggests reliance on memorized patterns, while low IDR indicates more robust, content-sensitive reasoning.
Terminology
Summary
The gist The translation into Arabic can suppress conventional contamination indicators while still allowing models to benefit from exposure, particularly those with stronger Arabic capabilities.
Data Contamination and Evaluation Validity
Data contamination undermines the validity of Large Language Model evaluation by enabling models to rely on memorized benchmark content rather than true generalization While prior work has proposed contamination detection methods, these approaches are largely limited to English benchmarks, leaving multilingual contamination poorly understood. Beyond this technical challenge, contamination has broader implications for reproducibility, fairness, and the trustworthiness of machine learning as a scientific discipline [Balloccu et al., 2024; Chen et al., 2025a].
Investigation into Multilingual Contamination Dynamics
In this work, we investigate contamination from a multilingual perspective. Specifically, we ask whether translating benchmarks into a low resources language,in our case, Arabic can act as a natural barrier to contamination or whether translation merely conceals memorization effects. We finetune open-weight models on different benchmarks with varying proportions of their Arabic-translated test set and evaluate their performance on the original English benchmark.
Contamination Detection Methods
To detect memorization, we extend the Tested Slot Guessing method with a choice-reordering strategy and incorporate Min-K% probability analysis, capturing both behavioral and distributional contamination signals. Our results show that translation into Arabic suppresses conventional contamination indicators, yet models still benefit from exposure to contaminated data, particularly those with stronger Arabic capabilities. To address this blind spot, we propose Translation-Aware Contamination Detection, which identifies contamination by comparing signals across multiple translated benchmark variants rather than English alone.
Translation-Aware Contamination Detection (TACD)
We introduce Translation-Aware Contamination Detection (TACD), a diagnostic procedure for identifying contamination-consistent behavior in multilingual evaluation settings. TACD is motivated by the observation that translation perturbs surface form while largely preserving semantic representations, allowing memorized benchmark content to remain exploitable even when English-only contamination checks fail. The protocol constructs multiple evaluation views for each instance x, denoted by x = (q, ci K i=1, y), where q is the question, ci are the answer choices, and y is the correct answer index.
TACD Signals and Interpretation
TACD computes two complementary signals: Index Recall Rate (IDR) and Cross-Lingual Consistency (CLC). Elevated IDR indicates reliance on memorized index-level associations rather than content-sensitive reasoning. Increasing CLC with higher poisoned exposure suggests growing representational invariance across translations without manifesting as explicit index recall—an effect that would not be detectable through IDR alone.
Empirical Findings Across Model Families
We evaluate TACD across three open-weight model families: Llama-3.2-1B-Instruct, Gemma-3-1B-it, and Qwen3-1.7B, under increasing levels of poisoned exposure (0%, 10%, 50%, and 100%). LLaMA exhibits low IDR across all poisoning levels, with values remaining below the random baseline and cross-lingual consistency increasing gradually with poisoned exposure. Qwen displays perfect cross-lingual consistency across all poisoning levels while maintaining IDR near chance, suggesting a model- and prompt-specific collapse in which predictions are weakly conditioned on input content. Gemma presents a distinct regime: while IDR remains near the random baseline, cross-lingual consistency increases monotonically with poisoning.
Implications for Evaluation Practice
These results suggest two implications for future evaluation practice. First, translation should not be treated as decontamination; multilingual benchmarks require detection pipelines that explicitly compare signals across translated variants. Second, contamination auditing should pair aggregate metrics with behavioral probes that are robust to surface-form changes. More broadly, as evaluation becomes increasingly multilingual, contamination-aware benchmarking must move beyond English-only assumptions to preserve fairness, transparency, and reproducibility.
A Dataset Complexity and Descriptive Statistics
MMLU is substantially larger than XQUAD and is the only dataset with explicit subject labels, whereas XQUAD provides extractive question–answer pairs without subject annotations. Lexical overlap between contexts and questions in XQuAD is modest (mean ≈8%, with upper-tail percentiles below 25%), indicating that questions are not simple rephrasings of their contexts. Answer spans are concise overall, with low median lengths and limited upper tails, consistent with XQuAD’s design toward short extractive answers. Figure 4 illustrates how closely Arabic→English translations align with the original English prompts in embedding space across MMLU subjects. The concentration of flow from highsimilarity bins toward many subjects indicates strong semantic preservation under translation. Overall, the figure supports the claim that translation preserves meaning at the representation level, which can mask contamination signals that rely on surface-form differences.
Figure 2 illustrates vocabulary size and type–token ratio (TTR) for MMLU and XQUAD.
Figure 3 shows context–question lexical overlap and answer length statistics for XQuAD.
Figure 4 shows flow diagram mapping Arabic→English translated items (left) to MMLU subject labels (right).
Table 1 shows results of English MMLU and XQuAD using the Evaluation Harness.
Table 2 shows TS-Guessing results on MMLU (MCQ) and XQuAD (QA) at different contamination levels.
Table 3 shows AUROC results of Mink++ on MMLU and XQuAD across models and contamination levels.
Table 4 reports the Index Recall Rate (IDR) and cross-lingual Algorithm 1 Translation-Aware Contamination Detection (TACD) signals across model families and poisoning levels.
References
[Balloccu et al., 2024] Simone Balloccu, Patr´ıcia Schmidtova, Mateusz Lango, and Ond ´ ˇrej Dusek. Leak, cheat, re-peat: Data contamination and evaluation malpractices in closed-source llms. Proceedings of the European Chapter of the Association for Computational Linguistics, 2024
[Brown et al., 2020] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020
[Chen et al., 2025a] Simin Chen, Yiming Chen, Zexin Li, et al. Benchmarking large language models under data contamination: A survey from static to dynamic evaluation. arXiv preprint arXiv:2502.
Improvements for AI systems
-
System can implement Translation-Aware Contamination Detection (TACD) to identify contamination by
comparing signals across multiple translated benchmark variants rather than English alone,
whichreliably exposes contamination even when English-only methods fail.
-
The improved system will utilize TS-Guessing and MinK++ alongside TACD to provide a
joint behavior of IDR and CLC
to characterize contamination-consistent behavior, distinguishing it from reasoning-driven consistency. -
The system can differentiate between model regimes by analyzing the interaction of signals; for instance, identifying when
IDR remains near the random baseline
whilecross-lingual consistency increases monotonically with poisoning,
suggestingrepresentational invariance across translations without manifesting as explicit index recall.
-
The system will adopt a dual evaluation strategy: using aggregate metrics like accuracy alongside behavioral probes that are
robust to surface form changes,
ensuring that contamination auditing moves beyond English-only assumptions. -
The system can incorporate model-specific interpretations, such as recognizing when a model exhibits
perfect cross-lingual consistency across all poisoning levels while maintaining IDR near chance,
highlightingmodel- and prompt-specific collapse in which predictions are weakly conditioned on input content.
Abstract
Data contamination can invalidate benchmark evaluation when a model benefits from memorized evaluation content rather than genuine generalization. Yet contamination is difficult to audit when the exposed content differs in language from the evaluation benchmark. We study this failure mode by deliberately exposing four open-weight instruction-tuned LLMs to Arabic translations of MMLU and XQuAD evaluation items at increasing exposure levels, then evaluating them on the original English tasks. This controlled setup is a proxy for contamination rather than a reconstruction of real-world pretraining leakage. We first test two English-centric post-hoc probes, TS-Guessing and Min-K%++, and find that their signals largely disappear under translated exposure: TS-Guessing remains weak except for model-specific positional recall on MMLU, while Min-K%++ stays at or below chance. At the same time, English MMLU performance increases with Arabic exposure, showing that the absence of an English contamination signal does not imply the absence of an exposure effect. We then introduce Translation-Aware Contamination Detection (TACD), a training-data-free diagnostic based on cross-lingual prediction consistency and choice reordering. Cross-lingual consistency is substantially higher than an independence baseline and generally increases relative to the clean condition, although its magnitude is model-dependent and not strictly monotonic. These results show that translation can conceal contamination-related effects from English-only probes and motivate multilingual diagnostics that are explicitly framed as evidence of contamination-consistent behavior rather than definitive membership tests.
Sources
- Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
- PaLM: Scaling Language Modeling with Pathways
- ConStat: Performance-Based Contamination Detection in Large Language Models
- Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
- Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
- OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
- LLaMA: Open and Efficient Foundation Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering