Dialects of Translationese Shape Language Model Learning

arXiv:2602.16469 · cs.CL · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Dialects of Translationese Shape Language Model Learning".

Jane: The paper was written by Jenny Kunz from Department of Computer and Information Science, Linköping University and Linköping University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title Discussion: Tom: The title "Dialects of Translationese Shows How Lexical Diversity and Source-Target Syntactic Similarity Shape Learning" is incredibly rich, Jane, and I think it sets up a very clear research question about how we train AI. It isn't just asking if translation makes data worse; it’s asking *how* what makes the data unique influences the learning process.

Jane: Right, and the authors are using twenty-four typologically diverse source languages, which is such a broad range of influence, allowing them to test that specific idea of "dialects." It means they aren're not just comparing English translated into other languages; they're looking at the variety of fingerprints coming from many different linguistic backgrounds.

Lu: What really grabs me is the idea that these "dialects" are measurable, suggesting a structured way for patterns to emerge. It implies that we can categorize and anticipate how certain types of translationese will affect model behavior based on their source language structure.

Meng: From an engineering standpoint, knowing the source-target syntactic similarity is so critical because it gives us a practical metric to predict performance shifts before we even train the models. We can see which types of inputs will stress the model in predictable ways.

Lalam: I think this approach suggests that we are moving toward a more nuanced understanding of multilingual data, recognizing that it’s not just one single "translationese" but a collection of distinct linguistic influences based on where the original content came from.

Tom: It sounds like the title sets up a very systematic way to study how various factors in translation influence the whole picture.

Summary Discussion: Tom: Now, looking at their summary, it seems they’ve found that this training on translated data has a clear "flatten" effect on models, which is a really important concept to unpack. It's not just that the model learns less; it actually loses certain nuances in its ability to represent language.

Jane: And specifically, they mention an increase in perplexity when applying those models back to native English text, which is a direct measure of how much worse their performance is compared to standard AI baseline models.

Lu: I find the finding that the surface lexical diversity—the variety of words—is a stronger predictor of lower perplexity in low-data settings really fascinating. It suggests that richness at the word level helps stabilize learning when resources are scarce.

Meng: But, Meng notes, this benefit doesn's consistent with grammatical performance, which is a crucial distinction for practical AI deployment. We can get good vocabulary but without structural integrity, and that's not very helpful for downstream tasks like parsing or generation.

Lalam: This flattening effect also has implications for us in the AI community because if we are relying on translated data to scale up our training, we might be inadvertently sacrificing the complexity and richness of native language.

Tom: So, this suggests that while having a lot of diverse words helps when you have limited data, it doesn' not necessarily means your models are becoming more grammatically sound.

Improvements Discussion: Tom: The next section focuses on how to improve or at least understand the limitations of these findings, and the paper suggests that typological similarity is a huge factor once we have enough data. This is a massive shift from what happens in low-data scenarios.

Jane: It seems like when you move into that larger dataset, the inherent structure of the source language starts to become much more important than just how many words there are in the corpus.

Lu: I love that we’re seeing this distinction between the 100MB and 1000MB models; it confirms that scaling up is what allows typological structure to finally assert its influence on model performance.

Meng: From an engineering standpoint, this means if we need high quality, reliable grammatical performance in a specific domain, we should prioritize sourcing data from a language that structurally aligns closely with our target language.

Lalam: That alignment—that structural similarity—is where the dialects of translationese start to become more predictable and beneficial for our LLMs because it ensures consistency across different training inputs.

Tom: It sounds like the key here is that while diversity helps when we are struggling with limited data, when we have a large amount of data, the structural relationship between two languages becomes a powerful driver of quality.

Conclusion: Tom: As we wrap up this discussion on "Training Models on Dialects of Translationese Shows How Lexical Diversity and Source-Target Syntactic Similarity Shape Learning," it’s clear that the source language really does matter when building these translated training corpora.

Jane: And it’s not just one factor, Lu, but a combination of how much diversity is present and how structurally similar the various parts of the team are to each other.

Lu: The most impactful vision for this research is that we are learning how to manage cultural and linguistic transfer intentionally, seeing these dialects not as errors but as structural characteristics that we can guide our own AI systems through.

Meng: I think the practical takeaway is that if an application requires strong grammatical competence, we absolutely cannot ignore the source language; the structure of the data must match our needs.

Lalam: For us in LLMs, it means we can achieve a deeper understanding of cultural nuance by being aware of these structural influences, allowing us to better serve diverse global populations.

Tom: So, when we use translated data for AI training, we are essentially shaping the learning process based on both the content and the final thoughts on this paper's implications.

Jane: It’s definitely a conversation that will continue as we look at even more large-scale models and diverse language sets.

Tom: We'll be back next week with another exciting breakthrough in AI research, so thank you all for joining us!

Jenny Kunz

Department of Computer and Information Science, Linköping University · Linköping University

cs.CL

Submitted: 2026-08-24

Updated: 2026-08-25

Comments: To appear at the Findings of EMNLP 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 84/100

The gist: This paper investigates how training on machine-translated data, a phenomenon known as "translationese," affects the linguistic development of small English language models.

Key concepts

Translationese
Refers to the language learned by AI models when trained extensively on translated data. The discussion explores how this type of training influences model performance and linguistic representation.
Lexical Diversity
The variety of words present in a corpus. Hosts discuss that having high lexical diversity can help stabilize learning, particularly when training data is scarce or limited.
Source-Target Syntactic Similarity
A metric measuring how structurally similar the source language and the target language are. This similarity is critical for predicting and ensuring strong grammatical performance in AI models.
Perplexity
A direct measure of how well an AI model performs compared to a standard baseline. An increase in perplexity when using translated data suggests a decline in the model's overall performance.

Terminology

Summary

This paper investigates how training on machine-translated data, a phenomenon known as translationese, affects the linguistic development of small English language models. As machine-translated text is increasingly used in multilingual NLP to compensate for scarce native data, understanding how the specific characteristics of different source languages shape a model's grammatical and lexical knowledge is critical for ensuring the quality and naturalness of multilingual AI systems.

The core research objective

The study aims to systematically analyze how translationese from 24 typologically and resourcediverse source languages shapes linguistic acceptability judgments and language modelling for different domains. By training small 125M-parameter models on English text translated from these diverse languages, the researchers can observe how the specific properties of a source language influence what a model learns. The study focuses on two primary dimensions of performance:

** General perplexity (language modeling performance) across various domains like Wikibooks and FineWeb. **

** Intrinsic linguistic acceptability (grammatical competence) using the BLiMP benchmark to test syntactic and morphological phenomena. **

Experimental methodology

The researchers utilized the Goldfish setup, training small GPT-style Transformer models on controlled data budgets of either 100MB or 1000MB. To ensure comparability, they used a fixed tokenizer and evaluated the models using high-quality English benchmarks. The source languages were selected for their diversity, spanning several families including Indo-European (Germanic, Slavic, Indo-Iranian), Uralic, Semitic, Dravidian, Austronesian, and Niger-Congo/Bantu. To relate model behavior to typological distance, the study computed cosine similarity over WALS syntactic features using lang2vec.

Key findings on performance

The results demonstrate that translationese has a flattening effect on the resulting models, which manifests as increased perplexity on native text and reduced linguistic acceptability scores compared to models trained on native English. The drivers of performance vary depending on the scale of data and the metric used:

** For general perplexity, surface lexical diversity in low-data settings (100MB) correlates with lower perplexity, but this relationship disappears at larger scales (1000MB). **

** For grammatical performance, typological similarity between the source language and English significantly improves grammatical performance given enough data. At the 1000MB scale, syntactic similarity showed a strong and significant positive correlation with overall BLiMP accuracy. **

** Cross-evaluation experiments revealed that models trained on translations from one language generalize better to translations from another if the source languages are typologically similar, suggesting that similar languages result in similar translations into English, i.e., in similar dialects of translationese. **

Implications for model training

The paper concludes that the choice of source language is a decisive factor when constructing translated training corpora. While lexical diversity may help in low-data regimes, structural similarity strongly and significantly predicts grammatical performance as data scales increase. Consequently, when using translated data to support grammatical generalization, selecting a structurally similar source language is highly beneficial. Ultimately, the research highlights that the source language of translated training data shapes learned linguistic knowledge in systematic ways.

Improvements for AI systems

Based on the empirical findings of this paper, I propose the following specific architectural and data-curation improvements for Large Language Models (LLMs) to enhance linguistic competence and cross-lingual transfer:

  1. Implement a Typological-Aware Data Mixture for Multilingual Pre-training.

The improved AI system will prioritize training data from source languages that share high syntactic similarity (WALS features) with the target language when the goal is grammatical accuracy. Specifically, when fine-tuning or pre-training an English model to improve its linguistic acceptability (BLiMP scores), the system will weight data from typologically similar sources (e.g., Germanic or Slavic) more heavily than lexically diverse but structurally distant sources to maximize the acquisition of complex dependencies like Agreement and Filler-Gap structures.

  1. Dynamic Data Scaling based on Lexical Diversity (TTR) Thresholds.

The improved AI system will utilize a two-stage training curriculum. In low-data regimes (<100MB), the system will optimize for high Type-Token Ratio (TTR) and Bigram TTR to minimize perplexity. As the data scale increases (>1000MB), the system will automatically shift its optimization objective from lexical richness to structural/syntactic similarity, as the paper proves that lexical diversity becomes a non-factor for performance at scale, while typological proximity becomes the dominant driver of grammatical competence.

  1. Syntactic-Similarity Guided Cross-Lingual Transfer.

The improved AI system will use Syntactic Proximity Mapping to select donor languages for zero-shot or few-shot transfer tasks. Instead of selecting donor languages based on lexical overlap or sheer corpus size, the system will select languages with high cosine similarity in their WALS syntactic feature vectors. This will allow the model to generalize better across dialects of translationese, enabling more robust performance on downstream tasks involving long-distance dependencies and morphosyntactic constraints.

  1. Translationese-Aware Evaluation and Mitigation.

The improved AI system will incorporate a Translationese Detection Layer during training to identify and potentially down-weight text that exhibits characteristic machine-translation artifacts (e.g., reduced morphological richness or flattened stylistic patterns). This prevents the model from developing impoverished linguistic preferences, ensuring that the model's output maintains high idiomaticity and naturalness on native-text benchmarks rather than just achieving low perplexity on translated test sets.

Abstract

Machine-translated data is widely used in multilingual NLP, particularly where native text is scarce. However, translated text differs systematically from native text. This phenomenon is known as translationese, and it reflects both traces of the source language and characteristic properties of translation itself. In this paper, we study how training on machine-translated data affects small English language models, focusing on how translationese from different source languages shapes linguistic acceptability judgments and language modeling for different domains. We train models on English text translated from 24 typologically and resource-diverse source languages, enabling a systematic analysis of how source language and corpus properties influence what models learn. Our results show that the source language has a clear impact on model behavior: general perplexity is more driven by the lexical diversity of the translated corpus, but grammatical performance is strongly correlated to typological similarity to English if trained on enough data. Even translation quality is a strong predictor of language modeling performance.

Sources

Related papers