GPT-Live-1 in the API: Full-Duplex Voice at $0.05/Min
Coffee Summary
- GPT-Live-1 became generally available in the OpenAI API on September 10, 2026.
- Full-duplex: listen and speak at once; handles interruptions, tone/pace/style via system prompt, noise/silence, telephony.
- Voice layer priced at $0.05 per minute, billed per second (not rounded up); backend models/tools bill separately.
- Can delegate reasoning/tools to a backend (e.g. GPT-6 Astra, Luna, third-party, or Codex).
- Speak eval CLAIM on launch post: ~80% fewer interruptions vs turn-based; Stoyanov quote cites 80% codebase simplification / 23K lines removed.
What happened
OpenAI launched GPT-Live-1 in the API on September 10, 2026, bringing ChatGPT’s natural, full-duplex voice conversations to developers. Model ID: `gpt-live-1`. Sessions use the Live endpoint (`v1/live/sessions` per model docs).
Unlike turn-based voice stacks, GPT-Live-1 can listen and speak at the same time, reasoning over incoming and outgoing audio together.
Why it matters
Classic voice agents chain STT → LLM → TTS. Each hop adds latency and fragile interruption logic. GPT-Live-1 collapses the voice layer into one full-duplex model, then optionally delegates deeper reasoning and tools to a backend text model or agent harness.
That split — natural voice front-end + swappable brain — is the architectural change builders should evaluate.
What changed
Capabilities called out by OpenAI
- **Interruption handling** in a single model (vs brittle cascade handoffs).
- **Tone, pace, and style** steered through the system prompt.
- **Background noise and silence** handling without narrating every step.
- **Long-session** context retention improvements (company claim).
- **Telephony** support for phone-call agents.
- Native **ASR transcripts** and **response text**.
- **Keyword biasing** and alphanumeric understanding.
- Native turn detection still available if you need explicit turn boundaries.
Pricing (verified from model docs + launch post)
| Component | Price / note |
| — | — |
| GPT-Live-1 voice layer | $0.05 per minute, billed per second (not rounded up to the next minute) |
| Backend model / tools | Billed separately at normal model/tool rates |
| Example backends named by OpenAI | GPT-6 Astra; Luna for high-volume tasks; third-party models; Codex pairing shown in launch materials |
Company-reported outcomes (CLAIMs)
- **Speak** early eval: GPT-Live-1 cut interruptions by **almost 80%** versus previous turn-based systems (OpenAI launch post).
- **Tony Stoyanov** (Co-Founder & CTO; quoted on the launch post): compared to a cascaded build, GPT-Live-1 **simplified the codebase by 80%** and **removed 23K lines of code**, enabling real-time patient conversations.
These are not independent benchmarks — reproduce on your traffic.
Builder checklist
1. Stand up a Live session with `gpt-live-1`; confirm audio duplex and interrupt behavior.
2. Add system-prompt style controls (tone/pace) before wiring tools.
3. Decide delegation: keep Live thin, send hard work to Astra / Luna / your agent.
4. Capture ASR transcripts + response text for logging/QA.
5. Load-test telephony path if you ship phone support.
6. Cost model: `$0.05/min` voice plus backend tokens/tools — measure both.
7. Compare against your current STT–LLM–TTS cascade on interruption rate and lines of glue code.
Who should care
- Product teams with voice UX (support, scheduling, tutoring, healthcare workflows).
- Platforms drowning in cascade glue who want a thinner voice layer.
- Anyone pricing realtime voice who needs per-second billing clarity.
Limitations
- Voice price is only the **front-end**; backend spend can dominate.
- Image/video modalities unsupported on `gpt-live-1` per model docs.
- Speak “~80% fewer interruptions” and Stoyanov “80% / 23K lines” are **launch-post CLAIMs**.
- Rate limits are concurrent-session based; Free tier unsupported (per model docs).
- Language/voice catalog is expanding over time — check current voice list rather than assuming parity with ChatGPT consumer voices.
What to do next
Prototype a 10-minute duplex dialog with interruption tests, then add one delegated tool path. Price a realistic call length at $0.05/min + backend. Only then decide whether to retire your STT–LLM–TTS cascade.
AIImpish Take
GPT-Live-1 is the cleanest OpenAI answer yet to “voice agents feel laggy and brittle.” Pay $0.05/min for the duplex mouth/ears, keep reasoning on a backend you control, and judge success on interruption rate and glue-code deleted — not on demo wow. If your cascade already works and cost is dominated by LLM tokens, migrate the voice layer first; don’t rewrite the whole agent on day one.
AIImpish