Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
| Source: arXiv AI
Tags: TTS, Thai, synthetic data, knowledge distillation, speech synthesis, low-resource NLP, open-source, on-device AI
Wayu-Paxa-TTS-Edge, an 82M-parameter Thai TTS model trained entirely on synthetic speech generated from a 15-second voice reference, achieves 85.5% of Gemini 3.1's keyword accuracy while enabling on-device inference — released open-source with its evaluation framework.
Details
Deploying TTS for low-resource languages typically requires either large voice-cloning models with costly inference or compact fixed-voice systems needing large speaker-specific corpora. This paper takes a third route: using a large voice-cloning model as a data generator to create a compact student from a 15-second voice reference, trained only on synthetic speech.\n\nThai adds specific challenges: ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. The team studies how text preparation, synthetic generation, quality filtering, rejection sampling, and frontend choices affect the resulting student model.\n\nThe 82M-parameter Wayu-Paxa-TTS-Edge achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1's score), 91.4% pause precision, 3.7% CER on Thai, and 1.1% CER on English. It outperforms its OmniVoice teacher model on pause placement and intra-word pause rates. The model and evaluation framework are open-sourced, lowering the barrier for Thai TTS development.