OpenAI has launched GPT-Live, a new ChatGPT capability enabling real-time voice and video conversations, according to the company's announcement. The feature is powered by an updated, faster version of GPT-4o optimized for low-latency, interactive use.
A Seamless AI Conversation
GPT-Live moves beyond turn-based voice assistant interactions. Per OpenAI's announcement, users can interrupt ChatGPT mid-response, and the model is designed to pause and adapt rather than finish a scripted reply.
Users can also stream live video from their device camera, letting ChatGPT visually analyze their surroundings. OpenAI's demonstrations show use cases such as feedback while practicing a presentation, step-by-step help with a math problem on a whiteboard, and plant identification with care tips.
How It Works
The underlying model is a version of GPT-4o optimized for this streaming format, processing audio and video as they arrive rather than waiting for a user to finish speaking or upload a clip. Key capabilities described by OpenAI include real-time audio/video streaming, natural interruption handling, multimodal understanding of simultaneous audio and video input, and low-latency response generation.
Availability
OpenAI says GPT-Live will roll out to ChatGPT Plus and Enterprise users over the coming weeks, starting on iOS and Android, with a desktop version to follow later.