CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment

| Source: arXiv AI

Tags: text-to-speech, CTC, LLM, streaming TTS, audio AI, INTERSPEECH

CTC-TTS replaces the standard MFA forced-aligner in LLM-based TTS with a CTC neural aligner and a bi-word interleaving strategy, achieving better streaming synthesis quality and lower latency — accepted at INTERSPEECH 2026.

Details

CTC-TTS addresses a practical bottleneck in streaming text-to-speech: accurate alignment between text and speech tokens. Current LLM-based TTS systems often rely on GMM-HMM alignment toolkits like the Montreal Forced Aligner (MFA), which are pipeline-heavy and brittle. CTC-TTS replaces MFA with a CTC-based neural aligner and introduces a bi-word interleaving strategy that better captures text-speech timing regularities. Two model variants are offered: CTC-TTS-L (token concatenation along sequence length) for higher audio quality, and CTC-TTS-F (embedding stacking along feature dimension) for lower latency. Both outperform fixed-ratio interleaving baselines and MFA-based approaches on both streaming synthesis and zero-shot tasks. The paper, accepted to INTERSPEECH 2026, is primarily a systems contribution — it solves a real engineering pain point in production TTS pipelines without requiring new training paradigms. The source is an arXiv revision with a minor typo fix.