Gradium AI Releases New Default TTS Model: 81.0% Hard-Case Pass Rate at 216 ms Time-to-First-Audio
| Source: MarkTechPost
Tags: Gradium AI, text-to-speech, TTS, voice AI, speech synthesis, ElevenLabs, Cartesia
Gradium AI replaced its default TTS model with one scoring 81.0% on a rigorous 500-sentence hard-case accuracy benchmark — beating Cartesia Sonic 3.6 (75.1%) and ElevenLabs v3 Conversational (65.4%) — while hitting 216 ms P50 latency with a 30 ms IQR, making it the front-runner for voice agents where accuracy on numbers, emails, and codes matters most.
Details
Gradium AI rolled out a new text-to-speech model as the default across its API and Studio on August 31, 2026. The swap is transparent — existing voice clones and integrations work unchanged, no migration required. The company benchmarked against four competitors on a 500-sentence hard-case eval set spanning five languages (EN, DE, FR, ES, PT) and 10 criteria targeting failure modes common in voice agents: acronyms, alphanumeric tokens, email addresses, dates, large numbers, and composite agent-turn scenarios (order lookups, IT tickets, insurance claims). Human raters scored strictly — one mispronounced digit fails the sentence. Results: Gradium 81.0%, Cartesia Sonic 3.6 75.1%, ElevenLabs v3 Conversational 65.4%, Fish Audio S2.1 Pro 49.5%, Inworld TTS 1.5 Max 46.5%. The eval set is open-sourced on Hugging Face under CC BY 4.0. On latency, Gradium reports 216 ms P50 on Coval's TTS benchmark — 170 ms faster than its predecessor — with a 30 ms IQR across 480 runs, the tightest spread of the five models tested. Inworld TTS 2 edges it on raw speed (166 ms), but Gradium's stated advantage is the best joint position: highest accuracy combined with sub-250 ms P50 latency. One caveat: Gradium designed and ran the benchmark itself. Publishing the eval set on Hugging Face allows independent reproduction, which will be the real test of these numbers.