ChatGPT’s new voice mode offers smoother conversations through the GPT-Live model, which listens and speaks at the same time, with three response levels that differ in speed and thinking time. The mode is available through the ChatGPT app on iPhone and Android devices and via web browsers, while the available model version varies between free and paid plans.
How to turn on voice mode
A conversation can be started by tapping the voice icon, which appears as a sound wave at the far right of the text box. The first time the mode is used, the app requests permission to access the device’s microphone. An animated floating circle then appears on the screen, allowing the user to speak naturally.
Voice mode remains active when the user switches to another app or locks the phone screen. Users can also interrupt ChatGPT or continue speaking without waiting for the response to finish, as the model is now less likely to interpret a brief pause as the end of a sentence, with the aim of making longer conversations flow more smoothly.
Three response levels
The voice model’s intelligence level can be changed using the settings icon at the top of the screen. GPT-Live starts automatically at the Instant level, while it can refer questions requiring research or deeper thinking to other models in the background. Selecting Medium or High may add a second or two to the response time, particularly when discussing complex topics repeatedly.
How has voice mode evolved?
The first version of Voice Mode, launched by OpenAI in 2023, relied on three separate systems: converting the user’s speech into text, generating a response using a language model, and then converting it into speech. The company later introduced Advanced Voice Mode using a multimodal model to speed up conversations and make them more natural, before GPT-Live became the default option in its place.
GPT-Live relies on an architecture that OpenAI describes as Full-Duplex, allowing the model to listen and speak at the same time. This helps reduce interruptions when the user pauses briefly and makes the exchange more like a natural conversation. One of the changes introduced by the model is the delivery of short vocal responses while the user is speaking, such as «Hmm» or «Yes».
According to the source material, GPT-Live can draw on more capable models in the background, such as GPT-5.5, when a question requires more thinking or research, while the voice conversation continues without a noticeable interruption.
Differences between plans and available features
Subscribers to ChatGPT Pro, Plus and Go receive the GPT-Live-1 model, while users on the free plan use GPT-Live-1 mini. According to the report, GPT-Live-1 mini does not allow users to customize the intelligence level, making this capability exclusive to paid plans.
GPT-Live’s fast response speed can be used for translation during live conversations. Users can also scroll down the screen to view a written transcript of what both parties are saying. If ChatGPT determines that an answer requires a visual element, it can create an interactive widget alongside the voice response to help clarify it.
GPT-Live does not currently support screen or video sharing. However, users can return to the previous voice model through ChatGPT’s settings by entering the Voice section and selecting a different model. The same section allows users to change ChatGPT’s voice and set the language, as well as configure Voice Mode as the default mode when the app starts.
