GPT-Live: OpenAI Bets the Interface Is Voice
GPT-Live combines full-duplex voice with backend model delegation. How live conversation, tool permissions, interruptions and application state fit together.
GPT-Live is OpenAI’s full-duplex voice interface: a model can listen and speak concurrently while delegating heavier work to a backend. The architectural idea is more useful than a rollout headline. Keep the conversation responsive while a separately configured system does the reasoning and tool work.
The current GPT-Live API documentation separates live audio interaction from backend execution. Listening and speaking can overlap. Your application still owns state, authorization and the meaning of an interruption.
Full-duplex is the actual product
A turn-oriented voice interface typically waits for an utterance boundary before producing the next answer. That interaction pattern does not, by itself, prove that the underlying audio transport is half-duplex.
Turn-oriented systems depend on decisions about when a speaker has finished. Short pauses, corrections and interruptions complicate those decisions. Full-duplex interaction changes the design space, but good behavior still has to be tested under real audio conditions.
Evaluate the entire interaction: capture, response timing, speech delivery, interruption handling and backend completion. Fast token generation alone does not make a conversation feel responsive. Test what happens when the user changes a request while a tool is already running.
| Concern | Turn-oriented design | Full-duplex application |
|---|---|---|
| Speech | Often organized around utterance boundaries | Can listen and speak concurrently |
| Interruption | Must recover interrupted conversation state | Must also coordinate active backend work |
| Authorization | Application must enforce permissions | Application must enforce permissions |
| Verification | Inspect tool results and accepted outcomes | Inspect tool results and accepted outcomes |
The clever part: fast mouth, slow brain
GPT-Live can delegate through the Responses API or through a client-controlled agent. The application selects the backend; the API design does not require one fixed model such as GPT-5.5. That allows the voice layer and the reasoning system to evolve independently.
Separating live conversation from backend work gives each component a clearer role. It also introduces coordination costs: task state, cancellation, duplicate tool requests and stale answers need explicit handling. This is a design option to evaluate, not proof that every monolithic alternative is inferior.
This is the part likely to outlast the product. A lot of AI is still built as if the whole game is picking the one best model. GPT-Live points at the more honest design: a small fast controller for presence and rhythm, a frontier model invoked only when the task earns it. The interface becomes an orchestrator.
Building with the API
The documented transports include WebRTC for browser interactions, WebSocket for server connections and SIP for telephony. Keep server credentials and privileged tool execution on the server. Voice-session duration, backend-model work and tools have separate billing implications; check the current pricing for the intended deployment.
Three takes
Why the launch matters
The text box made AI useful because it was simple, universal, and precise, and it is not going anywhere — real work still needs diffs, tables, citations, logs, things you can inspect. Voice will not replace that. But it can own the moments where you are moving faster than a keyboard, where the task is exploratory, or where the whole interaction depends on timing.
The useful promise is continuous conversation connected to capable software. The engineering obligation is preserving intent when speech overlaps with actions. An interruption does not necessarily cancel backend work: applications must decide how to stop, continue or safely reconcile it.
Sources: GPT-Live API guide; GPT-Live 1 model; Pricing. Documentation checked September 20, 2026.