Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation
cs.CL
Submitted: 2026-04-13
Updated: 2026-08-28
Comments: Accepted to EMNLP 2026 Main Track. Website is in https://ljvmiranda921.github.io/polyglot-teachers/
Code: https://github.com/bespokelabsai/curator
Project page: https://ljvmiranda921.github.io/polyglot-teachers
License: http://creativecommons.org/licenses/by/4.0/
The gist: Synthesizing supervised finetuning (SFT) data from language models (LMs) to teach smaller models multilingual tasks has become increasingly common.
Terminology
Abstract
Synthesizing supervised finetuning (SFT) data from language models (LMs) to teach smaller models multilingual tasks has become increasingly common. However, teacher model selection is often ad hoc, typically defaulting to the largest available option, even though such models may have significant capability gaps in non-English languages. This practice can result in poor-quality synthetic data and suboptimal student downstream performance. In this work, we systematically characterize what makes an effective multilingual teacher. We combine intrinsic measures of data quality with extrinsic student model performance in a metric we call Polyglot Score. We evaluate 10 LMs across 6 typologically diverse languages, generating over 1.4M SFT examples and training 240 student models. Our analyses reveal that model scale alone does not significantly predict teacher effectiveness: the most effective teachers we identify are consistently smaller than the largest models evaluated, and their ranking is stable across student base model families. Instead, data qualities such as prompt diversity, length, and response fluency capture 93.3% of the variance in intrinsic data quality and predict student performance. Finally, we provide practical recommendations, including matching the model families of teacher-student pairs and generating responses to existing prompts or translating them from English, which can yield improvements for less-resourced languages. We hope that our work advances data-centric research in multilingual synthetic data and LM development.
Sources
- Aya 23: Open Weight Releases to Further Multilingual Progress
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
- On the Diversity of Synthetic Data and its Impact on Training Large Language Models
- Training Verifiers to Solve Math Word Problems
- Command A: An Enterprise-Ready Large Language Model
- Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
- FastText.zip: Compressing text classification models
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- Scaling Laws for Neural Language Models
- EuroLLM-9B: Technical Report
- EuroLLM: Multilingual Language Models for Europe
- Towards Resource-Efficient LLMs: End-to-End Energy Accounting of Distillation Pipelines
- Bactrian-X: Multilingual Replicable Instruction-Following Models with Low-Rank Adaptation
- No Language Left Behind: Scaling Human-Centered Machine Translation
- GPT-4o System Card
- Tiny Aya: Bridging Scale and Multilingual Depth
- Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
- BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering