Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech
| Source: arXiv AI
Tags: ASR, code-switching, multilingual, speech recognition, Yoruba, audio LLM
A rigorous evaluation of 11 ASR and audio LM systems on English-Yoruba code-switched speech finds that aggregate WER masks critical failures: the top-WER ASR model is statistically indistinguishable from an audio LM on WER but far worse on every switch-localized metric.
Details
Code-switching — mixing two languages mid-utterance — is common in multilingual communities but poorly characterized in ASR benchmarks that report only aggregate WER. This paper evaluates six ASR models and five audio LMs on a deterministic 2000-utterance English-Yoruba evaluation set with switch-localized diagnostics. Beyond WER, the authors report switch entry token error rate (SETER), windowed switch point error rates, language-specific error rates, and diacritic-insensitive WER. The headline finding: the best-WER system (an ASR model) is statistically indistinguishable from a leading audio LM on aggregate WER, but the audio LM is significantly better on every switch-localized metric. Yoruba token recognition near-completely collapses for almost all systems (error rate 0.97), while English tokens are handled far better. Several generative audio LMs fail as faithful transcribers — they produce translations, verbose outputs, and prompt leakage with strong prompt dependence. The paper releases evaluation manifests, metric implementations, and scripts for reproducible benchmarking. This work is important for teams building ASR products for multilingual African markets and for researchers developing multilingual speech models.