Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
summary
The gist
The gist The translation into Arabic can suppress conventional contamination indicators while still allowing models to benefit from exposure, particularly those with stronger Arabic capabilities.
In short
The research investigated how translating benchmark data into Arabic affects data contamination detection in large language models. The study found that translation can suppress traditional contamination indicators, but models still benefit from exposure to contaminated data, especially those with strong Arabic skills. This necessitates a new method, Translation-Aware Contamination Detection (TACD), which compares signals across translated versions.
Key concepts
- Data Contamination
- This occurs when a model relies on memorized answers from the training data rather than learning true general knowledge. It undermines the validity of model evaluations because it makes models appear better than they actually are, as they simply 'cheat' by recalling specific examples.
- Translation-Aware Contamination Detection (TACD)
- This is a diagnostic procedure designed to find contamination in multilingual settings. It works by creating multiple evaluation views from the same data instance—one for the original English and others for its Arabic translations—and comparing the resulting signals to spot memorization effects that might be hidden by translation.
- Cross-Lingual Consistency (CLC)
- CLC measures how consistent a model's behavior is when tested across different language versions of the same prompt. When CLC increases with poisoned exposure, it suggests that the model develops a general understanding of the concept rather than just memorizing specific content.
- Index Recall Rate (IDR)
- IDR tracks whether a model relies on recalling specific index-level associations or reasoning based on content. High IDR suggests reliance on memorized patterns, while low IDR indicates more robust, content-sensitive reasoning.
Terminology used across episodes
This episode discusses
- Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora · Paper Radio
- Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
- PaLM: Scaling Language Modeling with Pathways
- ConStat: Performance-Based Contamination Detection in Large Language Models
- Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
- Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
- OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
- LLaMA: Open and Efficient Foundation Language Models
The paper
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora · Read on arXiv
Department of Electrical and Computer Engineering, Maroun Semaan Faculty of Engineering and Architecture, American University of Beirut
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Obscuring Data Contamination Through Translation".
Jane: The gist The translation into Arabic can suppress conventional contamination indicators while still allowing models to benefit from exposure, particularly those with stronger Arabic capabilities.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the show, everyone. Today we’re talking about a paper that tackles how we check if large language models are cheating on benchmarks, specifically data contamination. We’ve got "Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora."
Jane: It sounds like they're looking at a problem where English-only tests might be missing something important when we look at models that handle other languages.
Lu: Exactly. The paper argues that translation into a low-resource language, like Arabic, can actually act as a barrier to contamination signals because it suppresses the usual indicators we look for in English benchmarks <ref:2601.14994#pg1>.
Meng: So the core idea is that if you translate the benchmark, you might hide the memorization effect, but it doesn't remove it entirely <ref:2601.14994#pg2>.
Tom: Right. It’s about finding a way to detect contamination that works across different languages without just relying on English tests <ref:2601.14994#pg3>.
Jane: The paper investigates this by fine-tuning open-weight models on Arabic datasets and then testing them against the original English benchmarks <ref:2601.14994#pg1>.
Lu: They extend a method called Tested Slot Guessing by adding a choice reordering strategy and Min-K percent probability analysis to catch both behavioral and distributional signals <ref:2601.14994#pg3>.
Meng: And their finding is that translation into Arabic suppresses those conventional contamination indicators, which is interesting because it suggests the language shift itself hides the pattern <ref:2601.14994#pg3>.
Lalam: But they found that models still benefit from being exposed to contaminated data, especially those with stronger Arabic capabilities <ref:2601.14994#pg1>.
Tom: That leads us into the second part of the paper—how they propose a way to see this problem more clearly. The authors introduce Translation-Aware Contamination Detection, or TACD <ref:2601.14994#pg3>.
Jane: So TACD is a diagnostic procedure that compares signals across multiple versions of a benchmark instead of just English ones <ref:2601.14994#pg3>.
Lu: They construct several evaluation views for each instance, which involves the question, the answer choices, and the correct answer index <ref:2601.14994#pg3>.
Meng: They use two signals to measure this: Index Recall Rate and Cross-Lingual Consistency <ref:2601.14994#pg3>.
Tom: Index Recall Rate tells us if a model is relying on memorized index-level associations instead of actually reasoning about the content <ref:2601.14994#pg3>.
Jane: And Cross-Lingual Consistency measures how much that consistency increases as the exposure to contaminated data grows across translations <ref:2601.14994#pg3>.
Lu: The empirical findings across models like Llama-three point two-1B-Instruct, Gemma-three-1B-it, and Qwen3 show different behaviors depending on the poisoning level <ref:2601.14994#pg4>.
Tom: For Llama, they saw low Index Recall Rate across all poisoning levels but cross-lingual consistency increasing gradually with the poisoned exposure <ref:2601.14994#pg4>.
Jane: Qwen showed perfect cross-lingual consistency regardless of the poisoning level while keeping its Index Recall Rate near chance, which points to a model and prompt specific collapse <ref:2601.14994#pg4>.
Meng: Whereas Gemma showed a different pattern where the Index Recall Rate stayed near random baseline but cross-lingual consistency increased steadily with poisoning <ref:2601.14994#pg4>.
Tom: So what does this mean for us practically? It suggests that we shouldn't just treat translation as some kind of decontamination process for evaluation pipelines <ref:2601.14994#pg5>.
Jane: Instead, it means we need detection pipelines that explicitly compare signals across translated versions rather than just sticking to English tests <ref:2601.14994#pg5>.
Lu: It also suggests that when auditing for contamination, we should pair aggregate metrics with behavioral probes that are robust to changes in the surface form of the text <ref:2601.14994#pg5>.
Meng: If we're going multilingual, we have to move past just assuming English-only checks are enough to maintain fairness and reproducibility <ref:2601.14994#pg5>.
Tom: So, looking at the full picture of "Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora," the authors show that translation into Arabic can suppress standard contamination cues, but models still show signs of being trained on contaminated data if you look at it through a multilingual lens <ref:2601.14994#pg1>.
Jane: The paper’s title really captures the main point: how translating data obscures the contamination signals that we usually look for in English benchmarks <ref:2601.14994#pg3>.
Lu: It’s a deep dive into why existing methods fail when models are exposed to diverse, multilingual data settings <ref:2601.14994#pg2>.
Meng: It gives us a concrete way to think about how memorization can hide itself depending on the language you're looking at <ref:2601.14994#pg3>.
Tom: So, the implication is that as evaluation gets more multilingual, we need smarter ways to detect contamination than just looking at one language’s statistics <ref:2601.14994#pg5>.
Jane: That means future benchmarking has to be much more careful about how it handles language differences in its testing protocols <ref:2601.14994#pg5>.
Conclusion: Tom: So, basically, this paper shows that translating those benchmark tests into Arabic actually hides the usual signs of data contamination we look for in English benchmarks <ref:2601.14994#pg3>.
Jane: That’s a big idea, Tom. It means if we only check for cheating using English tests, we might miss problems happening in other languages entirely <ref:2601.14994#pg5>.
Lu: Exactly. The authors are setting up this whole framework called Translation-Aware Contamination Detection to see those signals across different versions of the benchmark instead of just one language <ref:2601.14994#pg3>.
Meng: So, if a model is fine with Arabic text, it might be exploiting contamination in a way that's invisible when you only look at English results <ref:2601.14994#pg3>.
Lalam: From my side, the LLaMA model showed that even with Arabic translation, we still see that reliance on memorized patterns if we look at the Index Recall Rate <ref:2601.14994#pg3>.
Tom: And they found a way to measure that through these two signals, Index Recall Rate and Cross-Lingual Consistency, which gives us a much richer picture <ref:2601.14994#pg3>.
Jane: It really changes how we think about evaluation pipelines. We can't just treat translation as a way to clean up data; it's like putting a filter on the signal <ref:2601.14994#pg5>.
Lu: The implication is that when we build these benchmarks, they have to be designed with multilingual contamination in mind from the start <ref:2601.14994#pg5>.
Meng: So, for someone building AI systems, it suggests you need to check your contamination audits across all the languages you're testing in <ref:2601.14994#pg5>.
Lalam: It’s about making sure our models are fair across different linguistic capabilities rather than just one specific language <ref:2601.14994#pg3>.
Tom: Yeah, so the title itself, "Obscuring Data Contamination Through Translation," really sums up the main challenge they're addressing here <ref:2601.14994#pg3>.
Jane: It’s a reminder that as AI gets more global, our way of checking for honesty in those models needs to get more global too <ref:2601.14994#pg5>.
Lu: And this work provides the tools to start building those better, multilingual detection methods from scratch <ref:2601.14994#pg3>.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck