OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time
| Source: THE DECODER
Tags: OpenAI, GPT-Live-1, voice AI, full-duplex, speech API, developer tools, conversational AI
OpenAI's GPT-Live-1 API brings full-duplex speech — simultaneous listening and speaking — to developers at $0.05/minute. Compared to its predecessor, interactivity scores jump from 45.4% to 80.1%, turn-taking latency drops from 1.4s to 0.8s, and tool-calling accuracy rises from 60% to 87%.
Details
OpenAI is opening GPT-Live-1 to developers via API. The model supports full-duplex speech — it can listen and speak simultaneously rather than waiting for the user to finish — enabling natural conversation flow in voice-based applications. It was already powering ChatGPT's Advanced Voice Mode before this API release. Benchmark improvements over the previous GPT-Realtime-2.1 are substantial across every measured dimension. Full-duplex interactivity jumps from 45.4% to 80.1%. Turn-taking latency drops from 1.4 seconds to 0.8 seconds, approaching natural conversation pace. Tool-calling accuracy — critical for voice agents that need to trigger backend actions — rises from 60% to 87%. A banking voice support benchmark shows pass rates climbing from 12.4% to 32%. The model is designed to be paired with different backend reasoning models depending on workload, letting developers balance depth, speed, and cost. Yelp is already using GPT-Live-1 for phone-based restaurant reservations; CTO Alex Levy cited improved call handling. The API ships with 12 new voices spanning different accents, dialects, and languages, plus built-in ASR transcription and response text output. Pricing at $0.05 per minute is not trivial. A deployment handling 10,000 minutes per day costs $500 daily, making cost modeling important before scaling voice features to high-volume production workloads. Full API documentation is pending.