Text to Speech News and AI Updates
Follow Text to Speech developments across AI companies, labs, and open-source projects.
Latest Text to Speech news, research, benchmarks, product releases, and industry adoption updates.
Latest Articles
- Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions — Researchers introduce Iterative Tensor Network Transformations (ITNTs), a framework enabling nonlinear operations on compressed tensor network data structures — demonstrating applications from 3D reactive flow field analysis to solving Max-SAT instances over spaces of up to 2^70 configurations.
- Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas — Cartesia's Sonic-3.6 takes #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on the Provider Voice board and 1,123 on the stricter Controlled Voice board — delivering sub-90ms time-to-first-audio at $49/1M characters, exactly half the price of ElevenLabs Eleven v3.
- A survey of AI-generated voices and their detection — A comprehensive survey of AI voice generation and detection covers TTS and voice cloning through deepfake audio detection, highlighting recent voice cloning scams targeting businesses and political leaders as evidence that synthetic voice fraud has moved beyond theoretical concern.
- Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair — Kozuchi Agent resolves 374/500 SWE-bench Verified instances (74.8%) using locally hosted Qwen3.5-27B with no fine-tuning, and ranks first among open-weight submissions on Multi-SWE-bench Java — making it the strongest language-agnostic open-weight software repair agent published to date.
- Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer — A single zero-initialized linear layer added to a 5B text-to-audio-video model enables state-of-the-art voice cloning across 30 speakers — outperforming five TTS baselines on speaker similarity — and the audio-only path runs approximately 30x faster than full audio-video diffusion.
- Content Based Video Narration of Gameplay with Vision Language Models — A no-training system turns any gameplay recording into esports-style commentary using a VLM and TTS, with a 3x3 frame mosaic trick that cuts image payload 9x and context-conditioned prompting to suppress repetitive narration — with a fully local TTS option on Apple silicon.
- Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers — Apple's production TTS system for Siri Expressive Voices achieves 16x real-time audio synthesis on-device with just 21MB peak runtime memory, using a Diffusion Transformer architecture that decouples temporal and depth processing for constant-memory inference.
- Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings — Alibaba's Qwen-Audio-3.0-TTS-Plus leads Artificial Analysis' Speech Arena with an Elo of 1,236, narrowly edging SpeechifyAI's Simba 3.2 — but its 16 chars/sec generation speed trails Sonic 3.5 (120 chars/sec) by a huge margin, limiting its practical value for latency-sensitive workloads.
- Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture — Re-Sonance chains Whisper ASR, Qwen LLM, and CosyVoice TTS to reconstruct natural-sounding speech from dysarthric input in real time, demonstrating improved intelligibility and naturalness for mild-to-moderate dysarthria in Mandarin while acknowledging limitations with severe cases.
- SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection — SONAR sets state-of-the-art on ASVspoof 2021 and in-the-wild audio deepfake benchmarks by explicitly disentangling low-frequency content from high-frequency synthesis artifacts using spectral contrastive learning — and converges four times faster than strong baselines while remaining architecture-agnostic.