OpenAI has just announced a new pair of conversational voice models, GPT‑Live‑1 and GPT‑Live‑1 mini, billed as full‑duplex systems that can speak and listen at the same time. The claim is that this design lets users interrupt naturally, making conversation flow much like a human exchange.

The company is also replacing its current Advanced Voice Mode in ChatGPT with GPT‑Live‑1 mini by default. Paid subscribers will be able to switch to the larger GPT‑Live‑1 model if they want more capability. According to the press briefing, the previous voice pipeline stitched together a speech‑to‑text model, a large language model, and a text‑to‑speech model – a cascade that often introduced latency and errors. The new full‑duplex approach avoids that round‑trip bottleneck.

Key points from the announcement:

  • Natural interruptibility – because the model can listen while it talks, users can break in without waiting for a turn‑taking signal.
  • Live translation – the model can translate speech on the fly, though early demos showed a heavy American accent when translating into Hindi.
  • Access to newer GPT models – the voice mode now has direct access to the latest text models (e.g., GPT‑5.5) for search, reasoning, or agentic tasks while the conversation continues.
  • Longer conversation windows – product lead Atty Eleti reported having 30‑ to 40‑minute voice chats during walks, suggesting the system can maintain context over extended periods.

OpenAI frames voice as a potential primary interface for complex work. The team envisions users relying on voice to control code assistants, manage data pipelines, or interact with multi‑agent systems hands‑free. Rivals such as Apple and Amazon have also been updating their assistants to be more conversational, and startups like Sesame are launching AI helpers with similarly natural dialogue.

“Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long‑running agentic work,” Eleti said. “The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work.”

Caveats. The new mode still has rough edges. In the demo, the Hindi translation sounded unnatural, and the system didn’t specify which languages it optimises for. Safeguards are in place for age‑appropriate responses and resources if the conversation turns to sensitive topics, but real‑world reliability remains to be seen.

Primary sources

Takeaway. OpenAI’s full‑duplex voice models represent a significant step toward voice‑first computing, removing the artificial “listen‑then‑speak” barrier that has limited natural conversation. Whether the technology lives up to the promise of seamless, long‑duration interaction will depend on how well it handles diverse languages, context retention, and real‑world latency.

Want this in your inbox every morning? Subscribe to the SpaghettiStories newsletter. Some links may be affiliate links. If you’re buying hardware to run local models, this affiliate link helps keep the lights on.