What Is GPT-Live 1 in 2026?
GPT-Live 1 is OpenAI's voice model for real-time conversations, made generally available in the API on 10 September 2026. OpenAI describes it as "a full-duplex voice model" that "can listen and speak at the same time, and delegate reasoning and tool use to a backend agent". Voice sessions cost 0.05 US dollars per minute, billed per second, with backend model and tool usage charged separately.
The design choice that makes it interesting for businesses is the split. GPT-Live handles the conversation: listening, speaking, handling interruptions. A separate backend, which you choose, does the thinking and uses the tools. That means you can put a voice front on a text workflow you already trust, rather than rebuilding your business logic inside a voice model.
| Attribute | What OpenAI documents in 2026 |
|---|---|
| Status | Generally available in the API since 10 September 2026 |
| Model ID and endpoint | gpt-live-1 on v1/live/sessions |
| Price | 0.05 USD per minute, billed per second, "not rounded up to the next whole minute" |
| Billed separately | Backend model and tool usage at their normal rates |
| Modalities | Audio and text in, audio and text out; no image or video |
| Supported | Streaming and function calling; structured outputs and fine-tuning not supported |
| Connections | WebRTC for browsers, WebSockets for servers, telephony and SIP for phone calls |
| Partner integrations | LiveKit, Twilio, Telnyx and Daily/Pipecat |
How Does the Live and Backend Split Work?
OpenAI separates two roles. "GPT-Live handles conversation. It listens, speaks, and decides when to ask the backend for help." The backend "reasons, uses tools, and returns results for GPT-Live to communicate". OpenAI's example: a caller asks about an order and adds a detail while the backend is checking its status, and GPT-Live keeps talking and explains the result when it arrives.
You choose how the backend runs, once per session:
- Responses delegation: OpenAI runs a hosted Responses model as the backend and passes conversation context and results between it and GPT-Live. Your application still runs your own function tools.
- Client delegation: you connect your own agent, model or service, and control its context, execution and returned results yourself.
To change modes you start a new session. OpenAI's guidance on where to put instructions is practical: keep speaking style in the live model's prompt, and business rules and tool workflows in the backend prompt.
How Much Does a Voice Agent Actually Cost in 2026?
The 0.05 US dollar per-minute rate covers only the voice session. Every question the backend answers is billed at that model's token rates, and every tool call at the tool's rate. OpenAI is explicit: "Backend Responses calls use the normal pricing for the configured model and tools."
| Cost line | How it is billed | What drives it |
|---|---|---|
| Voice session | 0.05 USD per minute, per second | Call length |
| Backend model | Selected model's per-token API rates | How often and how deeply the backend reasons |
| Tools | Standard tool rates, for example web search at 10 USD per 1,000 calls plus content tokens | How many lookups each call needs |
A labelled illustration, using only the published session rate: a voice session of five minutes costs 0.25 US dollars before any backend work, and 1,000 such calls cost 250 US dollars in session time. The backend model can move the total far more than the session rate does. Pairing the voice layer with GPT-6 Luna for routine lookups rather than GPT-6 Astra is the single biggest cost lever. These figures are arithmetic on published rates at invented volumes, not observed spend.
How Many Calls Can Run at Once?
GPT-Live rate limits are counted in concurrent sessions, and the Free tier is not supported. On 6 October 2026 OpenAI simplified its API usage tiers from five to three, Build, Launch and Grow, which reach at 5, 100 and 500 US dollars of total credit purchases.
| API usage tier (from 6 October 2026) | Concurrent sessions |
|---|---|
| Build | 50 |
| Launch | 300 |
| Grow | 500 |
A new account on the Build tier can run 50 simultaneous voice sessions. That is plenty for a pilot and a hard ceiling for a campaign that expects call spikes.
Who Is Responsible for What in a Voice Agent?
OpenAI puts the consequential decisions with your application, not the model. "Your application checks permissions, obtains required confirmations, runs functions that access your systems, and saves task progress." Backend work can continue when a caller interrupts, and "your application decides whether to finish or cancel it". In both delegation modes, OpenAI says, "your application controls permissions and business records".
That is the right split, and it means the quality of a voice agent depends heavily on the confirmations your application insists on. A refund, a booking change or an address update should require an explicit confirmation step that your code enforces, not one the conversation merely implies.
A new option for routing spoken requests
On 6 October 2026 OpenAI released the Decisions API in beta, which returns typed answers such as a choice from a fixed list, a probability or a score. OpenAI documents using it with GPT-Live through client delegation "to choose actions from voice requests and report their results to the user". For a support line, that is a cheap way to route a caller's request to the right action before any heavier backend work starts. It is covered in our GPT-6 Luna guide.
What Should Indian Businesses Check Before Using Voice Agents in 2026?
Two things. First, outbound calling for commercial purposes in India now sits under TRAI's revised commercial communication rules. If a voice agent places outbound promotional or service calls, read our guide to TRAI's A2P call rules and the TCCCPR Third Amendment before you dial. Inbound support lines, where the customer calls you, are a different situation.
Second, languages. The GPT-Live 1 model page does not publish a list of supported languages. Test Hindi, regional languages and code-switched speech on real calls before committing a production line, and keep a human fallback for callers the agent cannot serve well.
How Should Marketing and Support Teams Use GPT-Live in 2026?
| Use | Fit in 2026 | Control to set |
|---|---|---|
| Inbound support line answering order and account questions | Strong. This is OpenAI's own example. | Read-only lookups; confirmations enforced in code for changes. |
| After-hours enquiry capture for a clinic or showroom | Strong. Short, bounded conversations. | Hand off to a person the next morning. |
| Lead qualification on an inbound callback | Good. | Log every answer to the CRM for review. |
| Outbound promotional calls in India | Only after checking TRAI's A2P rules. | Consent and pre-declaration first. |
| Payments or account changes by voice | Weak without strong confirmation. | Your application must enforce the confirmation. |
What Are the Common Mistakes With GPT-Live in 2026?
- Budgeting only the 0.05 USD session rate. Backend model and tool usage is billed on top.
- Putting business rules in the voice prompt. OpenAI advises keeping them in the backend.
- Letting the conversation stand in for confirmation. Your application must check permissions and obtain confirmations.
- Ignoring the concurrency ceiling. Build-tier accounts get 50 concurrent sessions.
- Assuming Indian language quality. No language list is published; test on real calls.
- Running outbound commercial calls in India without reading TRAI's rules.
Key Takeaways for 2026
- GPT-Live 1 became generally available in the API on 10 September 2026.
- Voice sessions cost 0.05 US dollars a minute, billed per second, with backend model and tool usage billed separately.
- GPT-Live handles the conversation and delegates reasoning to a backend you choose, through Responses delegation or client delegation.
- Concurrency runs from 50 sessions on the Build tier to 500 on Grow; the Free tier is not supported.
- Your application, not the model, owns permissions, confirmations and business records.
- In India, check TRAI's A2P rules before placing outbound commercial calls, and test language quality on real calls.
Distk helps growth teams across India and internationally design voice agents with the confirmations and fallbacks in the right place, pick the backend model that keeps a call affordable, and stay inside India's commercial calling rules. If a voice line is on your 2026 plan, that design is where we start.
Sources
- OpenAI developer changelog, entry for 10 September 2026.
- OpenAI, GPT-Live 1 model page.
- OpenAI, Getting started with GPT-Live.
- OpenAI, Voice agents guide.
- OpenAI API pricing.