GPT-Live-1 in the API: Full-Duplex Voice at $0.05/Min

Coffee Summary

  • GPT-Live-1 became generally available in the OpenAI API on September 10, 2026.
  • Full-duplex: listen and speak at once; handles interruptions, tone/pace/style via system prompt, noise/silence, telephony.
  • Voice layer priced at $0.05 per minute, billed per second (not rounded up); backend models/tools bill separately.
  • Can delegate reasoning/tools to a backend (e.g. GPT-6 Astra, Luna, third-party, or Codex).
  • Speak eval CLAIM on launch post: ~80% fewer interruptions vs turn-based; Stoyanov quote cites 80% codebase simplification / 23K lines removed.

What happened

OpenAI launched GPT-Live-1 in the API on September 10, 2026, bringing ChatGPT’s natural, full-duplex voice conversations to developers. Model ID: `gpt-live-1`. Sessions use the Live endpoint (`v1/live/sessions` per model docs).

Unlike turn-based voice stacks, GPT-Live-1 can listen and speak at the same time, reasoning over incoming and outgoing audio together.

Why it matters

Classic voice agents chain STT → LLM → TTS. Each hop adds latency and fragile interruption logic. GPT-Live-1 collapses the voice layer into one full-duplex model, then optionally delegates deeper reasoning and tools to a backend text model or agent harness.

That split — natural voice front-end + swappable brain — is the architectural change builders should evaluate.

What changed

Capabilities called out by OpenAI

  • **Interruption handling** in a single model (vs brittle cascade handoffs).
  • **Tone, pace, and style** steered through the system prompt.
  • **Background noise and silence** handling without narrating every step.
  • **Long-session** context retention improvements (company claim).
  • **Telephony** support for phone-call agents.
  • Native **ASR transcripts** and **response text**.
  • **Keyword biasing** and alphanumeric understanding.
  • Native turn detection still available if you need explicit turn boundaries.

Pricing (verified from model docs + launch post)

| Component | Price / note |

| — | — |

| GPT-Live-1 voice layer | $0.05 per minute, billed per second (not rounded up to the next minute) |

| Backend model / tools | Billed separately at normal model/tool rates |

| Example backends named by OpenAI | GPT-6 Astra; Luna for high-volume tasks; third-party models; Codex pairing shown in launch materials |

Company-reported outcomes (CLAIMs)

  • **Speak** early eval: GPT-Live-1 cut interruptions by **almost 80%** versus previous turn-based systems (OpenAI launch post).
  • **Tony Stoyanov** (Co-Founder & CTO; quoted on the launch post): compared to a cascaded build, GPT-Live-1 **simplified the codebase by 80%** and **removed 23K lines of code**, enabling real-time patient conversations.

These are not independent benchmarks — reproduce on your traffic.

Builder checklist

1. Stand up a Live session with `gpt-live-1`; confirm audio duplex and interrupt behavior.

2. Add system-prompt style controls (tone/pace) before wiring tools.

3. Decide delegation: keep Live thin, send hard work to Astra / Luna / your agent.

4. Capture ASR transcripts + response text for logging/QA.

5. Load-test telephony path if you ship phone support.

6. Cost model: `$0.05/min` voice plus backend tokens/tools — measure both.

7. Compare against your current STT–LLM–TTS cascade on interruption rate and lines of glue code.

Who should care

  • Product teams with voice UX (support, scheduling, tutoring, healthcare workflows).
  • Platforms drowning in cascade glue who want a thinner voice layer.
  • Anyone pricing realtime voice who needs per-second billing clarity.

Limitations

  • Voice price is only the **front-end**; backend spend can dominate.
  • Image/video modalities unsupported on `gpt-live-1` per model docs.
  • Speak “~80% fewer interruptions” and Stoyanov “80% / 23K lines” are **launch-post CLAIMs**.
  • Rate limits are concurrent-session based; Free tier unsupported (per model docs).
  • Language/voice catalog is expanding over time — check current voice list rather than assuming parity with ChatGPT consumer voices.

What to do next

Prototype a 10-minute duplex dialog with interruption tests, then add one delegated tool path. Price a realistic call length at $0.05/min + backend. Only then decide whether to retire your STT–LLM–TTS cascade.

AIImpish Take

GPT-Live-1 is the cleanest OpenAI answer yet to “voice agents feel laggy and brittle.” Pay $0.05/min for the duplex mouth/ears, keep reasoning on a backend you control, and judge success on interruption rate and glue-code deleted — not on demo wow. If your cascade already works and cost is dominated by LLM tokens, migrate the voice layer first; don’t rewrite the whole agent on day one.