Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions

| Source: Apple ML Research

Tags: Apple, multilingual, low-resource languages, knowledge transfer, NLP, pretraining

Apple researchers propose LINK, a pretraining data intervention that improves multilingual knowledge transfer by randomly swapping English words with target-language translations — requiring only a bilingual vocabulary and achieving up to 2x faster convergence for low-resource languages.

Details

Cross-lingual knowledge transfer is the mechanism by which multilingual models extend capabilities from high-resource languages (mainly English) to low-resource ones. Existing methods for improving this transfer typically require parallel corpora, translation systems, or auxiliary models — resources often unavailable for truly low-resource languages. Apple's LINK (Lexical INtervention for Knowledge transfer) takes a radically simpler approach. LINK works at the data preprocessing stage: randomly selected words in a portion of the high-resource English training corpus are replaced with their word-level translations using a bilingual vocabulary. No model architectural changes, no parallel data, no auxiliary training steps — just a bilingual word list, obtainable at near-zero cost for virtually any language. Authors: Anastasiia Sedova, Natalie Schluter, Skyler Seto, and Maartje ter Hoeve (Schluter, Seto, and ter Hoeve as equal contributors). Evaluation across 8 languages and 5 model sizes shows notable downstream task improvements, with up to 2x speedup in reaching equivalent target-language performance. The method's simplicity makes it immediately deployable by any team doing multilingual pretraining without significant infrastructure investment.