Scaling Laws for Mixture Pretraining Under Data Constraints
| Source: Apple ML Research
Tags: Apple, scaling laws, mixture pretraining, multilingual, low-resource languages, data efficiency
Apple ML Research finds that in mixture pretraining, scarce domain-specific data can be repeated 15–20 times before hitting diminishing returns, and introduces a repetition-aware scaling law to help practitioners compute optimal data mix configurations — validated across 2,000+ training runs.
Details
A key challenge in pretraining large language models is balancing scarce but valuable target-domain data against abundant generic data. Too little target data underexposes the model; too much causes overfitting through excessive repetition. Apple researchers Anastasiia Sedova, Skyler Seto, Natalie Schluter, and Pierre Ablin studied this trade-off across more than 2,000 training runs spanning multiple model sizes, data types including multilingual, domain-specific, and quality-filtered mixtures. Their central finding: repetition is the dominant driver of target-domain performance, and mixture training tolerates far higher repetition than single-source training. Scarce target corpora can be reused 15–20 times, with the optimal count depending on target data size, compute budget, and model scale. The paper introduces a repetition-aware mixture scaling law that accounts for the decreasing marginal value of repeated target tokens and the regularizing effect of generic data. The law provides a principled formula for computing mixture configurations before committing to expensive training runs — directly useful for teams training multilingual models or domain-specific variants where data is inherently limited.