Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains
| Source: arXiv AI
Tags: multilingual NLP, cross-lingual transfer, XLM-R, Gemma 2, EMNLP 2026, low-resource languages
A trainable adaptive sampling strategy for multilingual model realignment assigns higher sampling probability to languages contributing more to the realignment loss — yielding +0.67 average performance on XLM-R and +0.60 on Gemma 2 9B across tasks. Accepted at EMNLP 2026 Findings.
Details
Realignment — updating multilingual models with parallel sentence pairs — is a key technique for improving cross-lingual transfer, especially for extremely low-resource languages. Existing methods use uniform random sampling across languages, which may undersample poorly-aligned languages that most need attention. This paper proposes adaptive sampling with trainable per-language probabilities. Languages contributing more to the realignment loss are sampled more frequently in subsequent batches, and the distribution evolves throughout training via an inner-outer optimization loop. The overhead is small. Results show consistent improvements: +0.67 average performance across all tasks with XLM-R, and +0.60 with Gemma 2 9B, compared to uniform realignment baselines. Importantly, the paper reports that gains are distributed across languages rather than concentrated in a few — a practical improvement for multilingual systems that must serve many locales. The method is model-agnostic and was validated across two substantially different architectures (XLM-R and Gemma 2 9B), suggesting the approach generalizes. Code is publicly available.