Meta AI enhances machine learning, NLP, and computer vision capabilities.
GPT-Live API
About GPT-Live API
The GPT-Live API provides a low-latency WebSocket endpoint for developers to integrate real-time voice interactions into applications. It supports bidirectional audio streaming with intermediate events for turn-taking, tool use, and reasoning handoff. The API exposes two model variants: GPT-Live-1 for higher voice quality and GPT-Live-1 mini for faster, cost-effective responses. Developers can expect function calling and tool integration mid-conversation, as well as the ability to hand off to other models like GPT-5.5 during complex queries. The API includes built-in automatic speech recognition and text-to-speech capabilities, eliminating the need for separate services. Access is currently available through a waitlist, with a developer preview rolling out in batches. Pricing is expected to be based on audio minutes streamed, with tiered rates depending on the model used. The API targets applications requiring real-time voice interactions, such as voice-first support agents, interpreters, or accessibility tools.
Key features
- Bidirectional audio streaming (PCM in, synthesized audio out)
- WebSocket-based transport for real-time communication
- Function calling and tool use mid-conversation
- Model handoff to GPT-5.5 or other chat models
- Multilingual ASR and TTS with translation support
- Turn-taking and reasoning handoff events
- Two model variants (GPT-Live-1 and GPT-Live-1 mini)
- Priority access for paid tiers
Use cases
- Voice-first support agents with real-time lookup and escalation
- Real-time interpreters for travel, healthcare, or field services
- Hands-free coding copilots for accessibility-focused users
Pros
- Real-time bidirectional audio streaming with WebSocket transport
- Built-in ASR and TTS in a single pipeline
- Support for function calling and model handoff during conversations
- Access to two model variants (GPT-Live-1 and GPT-Live-1 mini)
- Designed for low-latency, duplex voice interactions
Cons
- No public general availability date confirmed
- Waitlist-only access for developers
- Pricing not yet officially published
- Expected to be 3–5× more expensive than text APIs per interaction
Frequently asked questions about GPT-Live API
What is the GPT-Live API?
The GPT-Live API is a low-latency WebSocket endpoint that enables real-time voice interactions in applications. It supports bidirectional audio streaming with intermediate events for turn-taking, tool use, and reasoning handoff, and includes built-in automatic speech recognition and text-to-speech capabilities.
Who should use the GPT-Live API?
The API is designed for developers building voice-first applications such as support agents, interpreters, accessibility tools, or hands-free coding copilots. It suits use cases requiring real-time, duplex voice interactions with reasoning and tool integration.
How do I get access to the GPT-Live API?
Access is currently available through a waitlist. Developers can join by signing into their OpenAI account, visiting the waitlist page, confirming their use case and projected volume, and waiting for batch invites as OpenAI grants access in waves.
What pricing model does the GPT-Live API use?
The API is expected to use a usage-based pricing model tied to audio minutes streamed, with tiered rates depending on the model variant (GPT-Live-1 or GPT-Live-1 mini). Bundle discounts may apply for combining ASR, TTS, and reasoning in one call.
Does the GPT-Live API support tool use and model handoff during a conversation?
Yes, the API supports function calling and tool integration mid-conversation, as well as the ability to hand off to other models like GPT-5.5 during complex queries, similar to the ChatGPT app's voice mode.
What are the limitations or risks associated with the GPT-Live API?
Potential risks include latency spikes inherited from the underlying model, policy and safety considerations requiring clean handling of refusals, and cost at scale due to audio minutes adding up faster than text tokens.