Text to Speech News and AI Updates
Follow Text to Speech developments across AI companies, labs, and open-source projects.
Latest Text to Speech news, research, benchmarks, product releases, and industry adoption updates.
Latest Articles
- TokenMapper: A Step Toward Interoperable Speech Token Translation — TokenMapper enables direct token-to-token translation between structurally different speech tokenizers (single vs. multi-codebook), reducing end-to-end latency by 4.8-94.5% versus waveform bridging — accepted at AACL-IJCNLP 2026.
- Flexible and Interpretable Accent Distance Measurements — Cambridge researchers demonstrate that articulatory inversion representations combined with optimal transport measure accent distance interpretably from any recording type—bridging the gap between phonetics research methods and the accent embeddings used in TTS systems.
- Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations — Neural controlled differential equations (CDEs) applied to TTS duration modeling produce continuous-time hidden states whose values evolve with phonetic content, improving emotion intensity tracking in synthesized speech — accepted at IEEE SLT 2026.
- Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds — Gradium (spun out of Paris-based Kyutai lab) launched Voice Design, letting developers generate entirely new synthetic voices from a text description alone — no reference audio, no speaker cloning. The feature is free across all plans and live in the API and Studio today.
- CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment — CTC-TTS replaces the standard MFA forced-aligner in LLM-based TTS with a CTC neural aligner and a bi-word interleaving strategy, achieving better streaming synthesis quality and lower latency — accepted at INTERSPEECH 2026.
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization — A two-stage speech anonymization framework replaces identifiable content via generative editing and applies flow-matching-based voice anonymization (F3-VA) — maintaining ASR, TTS, and SER model utility better than VoicePrivacy Challenge baselines while improving privacy protection.
- X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System — X-Translator is an open-source real-time speech-to-speech translation system that preserves speaker voice identity across languages—using streaming ASR, machine translation, and a session-level speaker prompt manager with online speaker diarization, evaluated against proprietary APIs on OpenSTBench.
- Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech — Wayu-Paxa-TTS-Edge, an 82M-parameter Thai TTS model trained entirely on synthetic speech generated from a 15-second voice reference, achieves 85.5% of Gemini 3.1's keyword accuracy while enabling on-device inference — released open-source with its evaluation framework.
- Traceable TTS: Toward Watermark-Free TTS with Strong Traceability — A watermark-free TTS traceability framework uses joint training with a discriminator to attribute synthetic speech to its source model without embedding watermarks — the first approach to achieve strong traceability while preserving (and slightly improving) audio quality.
- Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations — A voice-first AI coaching system deployed to 40,000+ managers at scale rehearses difficult workplace conversations, with an end-to-end speech model achieving 3x lower latency and 8x lower cost than a cascaded LLM approach, though the cascaded system won out for coaching quality in production.