Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

| Source: MarkTechPost

Tags: Gradium, Kyutai, voice AI, text-to-speech, synthetic voice, ElevenLabs

Gradium (spun out of Paris-based Kyutai lab) launched Voice Design, letting developers generate entirely new synthetic voices from a text description alone — no reference audio, no speaker cloning. The feature is free across all plans and live in the API and Studio today.

Details

Gradium has shipped Voice Design, a feature that converts a written description into a deployable synthetic voice in 3–5 seconds. The input is a casting-brief-style prompt of 1–500 characters (English, French, Spanish, Portuguese, or German) specifying attributes like age, accent, pitch, pace, timbre, and intended use. One request returns 1–5 candidate voices; the chosen candidate is promoted to a production voice slot via a four-call API flow. The key differentiator from cloning: no source audio, no speaker consent, no licensing overhead. Gradium draws on the Kyutai research lineage and positions Voice Design as the answer for briefs that no catalog can satisfy — regional accents, niche personas, age-specific registers. In a blind pairwise benchmark across 7,627 comparisons and six publicly accessible voice-design systems in five languages, Gradium reported a 72.6% win rate, 13.6 percentage points ahead of ElevenLabs' eleven_ttv_v3 at 59.0%, followed by Inworld at 44.8%. The benchmark was conducted on accent prompts by native-speaker evaluators. Kept voices use the same streaming TTS endpoint as catalog voices, with no added latency. The free tier holds 5 custom voice slots shared with clones; paid plans offer 1,000. Candidates expire after 30 days if not converted; converting is free.