Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR
| Source: Apple ML Research
Tags: Apple, speech recognition, code-switching, Mandarin, ASR, pseudo-labeling, multilingual
Apple researchers apply iterative pseudo-labeling to Mandarin-English code-switching speech recognition for the first time, achieving 6.35% and 8.29% Mix Error Rate reductions on SEAME benchmarks by bootstrapping from unlabeled audio.
Details
Code-switching ASR — where speakers alternate languages mid-utterance, common in Mandarin-English conversations across Southeast Asia — is notoriously difficult due to limited labeled training data. Authors Qu Yang, Cakra Wardhana, and Tim Ng at Apple apply iterative pseudo-labeling to this problem, which they describe as a first in the literature. The approach runs in three phases: generating pseudo-labels from a large unlabeled corpus to create a semi-supervised dataset, training the model in two stages (pretraining then fine-tuning on supervised code-switching data), and iterative refinement passes that progressively improve labeling accuracy. The method reduces dependence on expensive labeled code-switching recordings. On the SEAME dataset, the approach achieved 6.35% Mix Error Rate (MER) reduction on the devman subset and 8.29% on devsge — meaningful gains in a task where marginal improvements are difficult. The two-stage training framework is likely generalizable to other code-switching language pairs. Results are directly applicable to voice product teams building for multilingual markets.