Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains
cs.CL, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Accepted to ENMLP 2026 Findings
Code: https://github.com/felixgaschi/multilingual-alignment-and-transfer
License: http://creativecommons.org/licenses/by/4.0/
The gist: Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs).
Terminology
Abstract
Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs). However, existing realignment methods rely on uniform and random sampling of parallel sentences across languages, which may be suboptimal under limited batch sizes. In practice, models may benefit from seeing certain languages more frequently, especially those that are poorly aligned, and the optimal distribution can evolve throughout training. In this work, we propose a simple yet effective adaptive sampling strategy that assigns trainable sampling probabilities to each language. Languages that contribute more to the realignment loss are sampled more frequently in subsequent batches, and the optimal distribution can evolve throughout training. Our method employs an inner-outer optimization loop with a small overhead, leading to consistent performance improvements and, more importantly, distributing the gains across languages. We observed a +0.67 average performance increase on all tasks with XLM-R, and +0.60 with Gemma 2 9B compared with uniform realignment. Furthermore, our method is robust across different models. Code available at https://github.com/felixgaschi/multilingual-alignment-and-transfer.
Sources
- Efficient Online Data Mixing For Language Model Pre-Training
- R3: Robust Rubric-Agnostic Reward Models
- Multilingual Alignment of Contextual Word Representations
- XNLI: Evaluating Cross-lingual Sentence Representations
- No Language Left Behind: Scaling Human-Centered Machine Translation
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- The Llama 3 Herd of Models
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- Myanmar XNLI: Building a Dataset and Exploring Low-resource Approaches to Natural Language Inference with Myanmar
- On First-Order Meta-Learning Algorithms
- Learning to Reweight Examples for Robust Deep Learning
- Gemma 2: Improving Open Language Models at a Practical Size
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering