Dialects of Translationese Shape Language Model Learning
summary
The gist
This paper investigates how training on machine-translated data, a phenomenon known as "translationese," affects the linguistic development of small English language models.
In short
The episode discusses a paper analyzing how training AI models on translated data affects learning. Hosts discuss that while high lexical diversity helps with limited data, grammatical performance relies more heavily on the structural similarity between source and target languages, suggesting careful data sourcing is crucial for quality.
Key concepts
- Translationese
- Refers to the language learned by AI models when trained extensively on translated data. The discussion explores how this type of training influences model performance and linguistic representation.
- Lexical Diversity
- The variety of words present in a corpus. Hosts discuss that having high lexical diversity can help stabilize learning, particularly when training data is scarce or limited.
- Source-Target Syntactic Similarity
- A metric measuring how structurally similar the source language and the target language are. This similarity is critical for predicting and ensuring strong grammatical performance in AI models.
- Perplexity
- A direct measure of how well an AI model performs compared to a standard baseline. An increase in perplexity when using translated data suggests a decline in the model's overall performance.
Terminology used across episodes
This episode discusses
- Dialects of Translationese Shape Language Model Learning · Paper Radio
- Goldfish: Monolingual Language Models for 350 Languages
- Preferences for Idiomatic Language are Acquired Slowly -- and Forgotten Quickly: A Case Study on Swedish
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
The paper
Dialects of Translationese Shape Language Model Learning · Read on arXiv
Jenny Kunz
Department of Computer and Information Science, Linköping University · Linköping University
Machine-translated data is widely used in multilingual NLP, particularly where native text is scarce. However, translated text differs systematically from native text. This phenomenon is known as translationese, and it reflects both traces of the source language and characteristic properties of translation itself. In this paper, we study how training on machine-translated data affects small English language models, focusing on how translationese from different source languages shapes linguistic acceptability judgments and language modeling for different domains. We train models on English text translated from 24 typologically and resource-diverse source languages, enabling a systematic analysis of how source language and corpus properties influence what models learn. Our results show that the source language has a clear impact on model behavior: general perplexity is more driven by the lexical diversity of the translated corpus, but grammatical performance is strongly correlated to typological similarity to English if trained on enough data. Even translation quality is a strong predictor of language modeling performance.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Dialects of Translationese Shape Language Model Learning".
Jane: The paper was written by Jenny Kunz from Department of Computer and Information Science, Linköping University and Linköping University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title Discussion: Tom: The title "Dialects of Translationese Shows How Lexical Diversity and Source-Target Syntactic Similarity Shape Learning" is incredibly rich, Jane, and I think it sets up a very clear research question about how we train AI. It isn't just asking if translation makes data worse; it’s asking *how* what makes the data unique influences the learning process.
Jane: Right, and the authors are using twenty-four typologically diverse source languages, which is such a broad range of influence, allowing them to test that specific idea of "dialects." It means they aren're not just comparing English translated into other languages; they're looking at the variety of fingerprints coming from many different linguistic backgrounds.
Lu: What really grabs me is the idea that these "dialects" are measurable, suggesting a structured way for patterns to emerge. It implies that we can categorize and anticipate how certain types of translationese will affect model behavior based on their source language structure.
Meng: From an engineering standpoint, knowing the source-target syntactic similarity is so critical because it gives us a practical metric to predict performance shifts before we even train the models. We can see which types of inputs will stress the model in predictable ways.
Lalam: I think this approach suggests that we are moving toward a more nuanced understanding of multilingual data, recognizing that it’s not just one single "translationese" but a collection of distinct linguistic influences based on where the original content came from.
Tom: It sounds like the title sets up a very systematic way to study how various factors in translation influence the whole picture.
Summary Discussion: Tom: Now, looking at their summary, it seems they’ve found that this training on translated data has a clear "flatten" effect on models, which is a really important concept to unpack. It's not just that the model learns less; it actually loses certain nuances in its ability to represent language.
Jane: And specifically, they mention an increase in perplexity when applying those models back to native English text, which is a direct measure of how much worse their performance is compared to standard AI baseline models.
Lu: I find the finding that the surface lexical diversity—the variety of words—is a stronger predictor of lower perplexity in low-data settings really fascinating. It suggests that richness at the word level helps stabilize learning when resources are scarce.
Meng: But, Meng notes, this benefit doesn's consistent with grammatical performance, which is a crucial distinction for practical AI deployment. We can get good vocabulary but without structural integrity, and that's not very helpful for downstream tasks like parsing or generation.
Lalam: This flattening effect also has implications for us in the AI community because if we are relying on translated data to scale up our training, we might be inadvertently sacrificing the complexity and richness of native language.
Tom: So, this suggests that while having a lot of diverse words helps when you have limited data, it doesn' not necessarily means your models are becoming more grammatically sound.
Improvements Discussion: Tom: The next section focuses on how to improve or at least understand the limitations of these findings, and the paper suggests that typological similarity is a huge factor once we have enough data. This is a massive shift from what happens in low-data scenarios.
Jane: It seems like when you move into that larger dataset, the inherent structure of the source language starts to become much more important than just how many words there are in the corpus.
Lu: I love that we’re seeing this distinction between the 100MB and 1000MB models; it confirms that scaling up is what allows typological structure to finally assert its influence on model performance.
Meng: From an engineering standpoint, this means if we need high quality, reliable grammatical performance in a specific domain, we should prioritize sourcing data from a language that structurally aligns closely with our target language.
Lalam: That alignment—that structural similarity—is where the dialects of translationese start to become more predictable and beneficial for our LLMs because it ensures consistency across different training inputs.
Tom: It sounds like the key here is that while diversity helps when we are struggling with limited data, when we have a large amount of data, the structural relationship between two languages becomes a powerful driver of quality.
Conclusion: Tom: As we wrap up this discussion on "Training Models on Dialects of Translationese Shows How Lexical Diversity and Source-Target Syntactic Similarity Shape Learning," it’s clear that the source language really does matter when building these translated training corpora.
Jane: And it’s not just one factor, Lu, but a combination of how much diversity is present and how structurally similar the various parts of the team are to each other.
Lu: The most impactful vision for this research is that we are learning how to manage cultural and linguistic transfer intentionally, seeing these dialects not as errors but as structural characteristics that we can guide our own AI systems through.
Meng: I think the practical takeaway is that if an application requires strong grammatical competence, we absolutely cannot ignore the source language; the structure of the data must match our needs.
Lalam: For us in LLMs, it means we can achieve a deeper understanding of cultural nuance by being aware of these structural influences, allowing us to better serve diverse global populations.
Tom: So, when we use translated data for AI training, we are essentially shaping the learning process based on both the content and the final thoughts on this paper's implications.
Jane: It’s definitely a conversation that will continue as we look at even more large-scale models and diverse language sets.
Tom: We'll be back next week with another exciting breakthrough in AI research, so thank you all for joining us!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language