AI Model Guide

GPT-Live 1 Voice Is Generally Available at $0.05 a Minute: Voice Agents for Business in 2026

OpenAI's new voice model listens and speaks at the same time and hands the thinking to a backend you choose. The headline rate is cheap. What a call actually costs depends on the backend, and what it can safely do depends on your application.

Distk Editorial Oct 2026 9 min read

GPT-Live 1 is OpenAI's full-duplex voice model, generally available in the API since 10 September 2026 as gpt-live-1 on v1/live/sessions. It listens and speaks at the same time and delegates reasoning and tool use to a backend, either an OpenAI-hosted Responses model or your own agent. Voice sessions cost 0.05 US dollars a minute, billed per second, while backend model and tool usage is billed separately at normal rates, which usually moves the total more than the session rate. Concurrency ranges from 50 sessions on the Build tier to 500 on Grow under the three usage tiers OpenAI introduced on 6 October, and the Free tier is not supported. OpenAI places permissions, confirmations and business records with your application. Indian businesses placing outbound commercial calls should check TRAI's A2P rules first and test language quality on real calls.

What Is GPT-Live 1 in 2026?

GPT-Live 1 is OpenAI's voice model for real-time conversations, made generally available in the API on 10 September 2026. OpenAI describes it as "a full-duplex voice model" that "can listen and speak at the same time, and delegate reasoning and tool use to a backend agent". Voice sessions cost 0.05 US dollars per minute, billed per second, with backend model and tool usage charged separately.

The design choice that makes it interesting for businesses is the split. GPT-Live handles the conversation: listening, speaking, handling interruptions. A separate backend, which you choose, does the thinking and uses the tools. That means you can put a voice front on a text workflow you already trust, rather than rebuilding your business logic inside a voice model.

AttributeWhat OpenAI documents in 2026
StatusGenerally available in the API since 10 September 2026
Model ID and endpointgpt-live-1 on v1/live/sessions
Price0.05 USD per minute, billed per second, "not rounded up to the next whole minute"
Billed separatelyBackend model and tool usage at their normal rates
ModalitiesAudio and text in, audio and text out; no image or video
SupportedStreaming and function calling; structured outputs and fine-tuning not supported
ConnectionsWebRTC for browsers, WebSockets for servers, telephony and SIP for phone calls
Partner integrationsLiveKit, Twilio, Telnyx and Daily/Pipecat

How Does the Live and Backend Split Work?

OpenAI separates two roles. "GPT-Live handles conversation. It listens, speaks, and decides when to ask the backend for help." The backend "reasons, uses tools, and returns results for GPT-Live to communicate". OpenAI's example: a caller asks about an order and adds a detail while the backend is checking its status, and GPT-Live keeps talking and explains the result when it arrives.

You choose how the backend runs, once per session:

To change modes you start a new session. OpenAI's guidance on where to put instructions is practical: keep speaking style in the live model's prompt, and business rules and tool workflows in the backend prompt.

How Much Does a Voice Agent Actually Cost in 2026?

The 0.05 US dollar per-minute rate covers only the voice session. Every question the backend answers is billed at that model's token rates, and every tool call at the tool's rate. OpenAI is explicit: "Backend Responses calls use the normal pricing for the configured model and tools."

Cost lineHow it is billedWhat drives it
Voice session0.05 USD per minute, per secondCall length
Backend modelSelected model's per-token API ratesHow often and how deeply the backend reasons
ToolsStandard tool rates, for example web search at 10 USD per 1,000 calls plus content tokensHow many lookups each call needs

A labelled illustration, using only the published session rate: a voice session of five minutes costs 0.25 US dollars before any backend work, and 1,000 such calls cost 250 US dollars in session time. The backend model can move the total far more than the session rate does. Pairing the voice layer with GPT-6 Luna for routine lookups rather than GPT-6 Astra is the single biggest cost lever. These figures are arithmetic on published rates at invented volumes, not observed spend.

How Many Calls Can Run at Once?

GPT-Live rate limits are counted in concurrent sessions, and the Free tier is not supported. On 6 October 2026 OpenAI simplified its API usage tiers from five to three, Build, Launch and Grow, which reach at 5, 100 and 500 US dollars of total credit purchases.

API usage tier (from 6 October 2026)Concurrent sessions
Build50
Launch300
Grow500

A new account on the Build tier can run 50 simultaneous voice sessions. That is plenty for a pilot and a hard ceiling for a campaign that expects call spikes.

Who Is Responsible for What in a Voice Agent?

OpenAI puts the consequential decisions with your application, not the model. "Your application checks permissions, obtains required confirmations, runs functions that access your systems, and saves task progress." Backend work can continue when a caller interrupts, and "your application decides whether to finish or cancel it". In both delegation modes, OpenAI says, "your application controls permissions and business records".

That is the right split, and it means the quality of a voice agent depends heavily on the confirmations your application insists on. A refund, a booking change or an address update should require an explicit confirmation step that your code enforces, not one the conversation merely implies.

A new option for routing spoken requests

On 6 October 2026 OpenAI released the Decisions API in beta, which returns typed answers such as a choice from a fixed list, a probability or a score. OpenAI documents using it with GPT-Live through client delegation "to choose actions from voice requests and report their results to the user". For a support line, that is a cheap way to route a caller's request to the right action before any heavier backend work starts. It is covered in our GPT-6 Luna guide.

What Should Indian Businesses Check Before Using Voice Agents in 2026?

Two things. First, outbound calling for commercial purposes in India now sits under TRAI's revised commercial communication rules. If a voice agent places outbound promotional or service calls, read our guide to TRAI's A2P call rules and the TCCCPR Third Amendment before you dial. Inbound support lines, where the customer calls you, are a different situation.

Second, languages. The GPT-Live 1 model page does not publish a list of supported languages. Test Hindi, regional languages and code-switched speech on real calls before committing a production line, and keep a human fallback for callers the agent cannot serve well.

How Should Marketing and Support Teams Use GPT-Live in 2026?

UseFit in 2026Control to set
Inbound support line answering order and account questionsStrong. This is OpenAI's own example.Read-only lookups; confirmations enforced in code for changes.
After-hours enquiry capture for a clinic or showroomStrong. Short, bounded conversations.Hand off to a person the next morning.
Lead qualification on an inbound callbackGood.Log every answer to the CRM for review.
Outbound promotional calls in IndiaOnly after checking TRAI's A2P rules.Consent and pre-declaration first.
Payments or account changes by voiceWeak without strong confirmation.Your application must enforce the confirmation.

What Are the Common Mistakes With GPT-Live in 2026?

Key Takeaways for 2026

Distk helps growth teams across India and internationally design voice agents with the confirmations and fallbacks in the right place, pick the backend model that keeps a call affordable, and stay inside India's commercial calling rules. If a voice line is on your 2026 plan, that design is where we start.

Sources

GPT-Live 1 in 2026: FAQs

What is GPT-Live 1?

OpenAI's full-duplex voice model for real-time conversations, generally available in the API since 10 September 2026. It can listen and speak at the same time and delegates reasoning and tool use to a backend model or agent you choose.

How much does GPT-Live 1 cost?

Voice sessions cost 0.05 US dollars per minute, billed per second and not rounded up to the next minute. Backend model and tool usage is billed separately at the normal rates for the configured model and tools.

What is the difference between Responses delegation and client delegation?

With Responses delegation, OpenAI runs a hosted Responses model as the backend and passes context between it and GPT-Live. With client delegation, you connect your own agent or service and control its context and execution. You choose the mode when creating the session.

How many simultaneous calls can GPT-Live handle?

Limits are counted in concurrent sessions: 50 on Build, 300 on Launch and 500 on Grow, the three usage tiers OpenAI introduced on 6 October 2026. The Free tier is not supported.

Can GPT-Live connect to phone calls?

Yes. OpenAI documents telephony and SIP connections for phone calls, WebRTC for browsers and WebSockets for servers, plus partner integrations with LiveKit, Twilio, Telnyx and Daily/Pipecat.

Can Indian businesses use GPT-Live for outbound calls?

Outbound commercial calling in India is governed by TRAI's revised commercial communication rules, including the A2P call provisions in the TCCCPR Third Amendment. Check those rules and your consent records before placing outbound promotional or service calls with a voice agent.

Put the confirmations in your code, not the conversation

Distk helps growth teams design voice agents with enforced confirmations and human fallbacks, choose the backend model that keeps calls affordable, and stay inside India's commercial calling rules.

Start the conversation →