What Is OpenAI Ultrafast Mode in 2026?
Ultrafast is OpenAI's fastest API service tier. OpenAI calls it "the fastest service tier in the OpenAI API" and says to use it "when speed justifies the higher cost". It launched for GPT-6 Astra on 29 September 2026, and on 8 October 2026 OpenAI added it for GPT-6.1 Sol, which had been announced as coming "in the following days". You turn it on per request by setting service_tier to ultrafast in the Responses API.
The changelog describes the effect as reducing "the time between generated output tokens". It matters for one kind of work: agents that make many tool calls in quick succession while a person waits on the result. For anything that runs overnight or in a batch, it is the wrong tool, because on both models it costs six times the Standard rate.
| Attribute (October 2026) | GPT-6 Astra Ultrafast | GPT-6.1 Sol Ultrafast |
|---|---|---|
| Released | 29 September 2026 | 8 October 2026 |
| How to enable | gpt-6-astra with service_tier: "ultrafast" | gpt-6.1-sol with service_tier: "ultrafast" |
| Price per 1M tokens (up to 272K input) | 60 USD input, 300 USD output | 12 USD input, 60 USD output |
| Data residency | US data residency and global processing only | US and EU data residency, and global processing |
| Availability | All API users, subject to rate limits | All API users, subject to rate limits |
How Much Does Ultrafast Cost in 2026?
On both models Ultrafast is six times the Standard rate. For GPT-6 Astra that is 60 US dollars per million input tokens and 300 per million output. For GPT-6.1 Sol it is 12 and 60. OpenAI's GPT-6.1 Sol model page states it directly: "Ultrafast mode prices are 6x Standard."
| Tier (per 1M tokens, up to 272K input) | GPT-6 Astra input / output | GPT-6.1 Sol input / output |
|---|---|---|
| Batch / Flex | 5.00 / 25.00 USD | 1.00 / 5.00 USD |
| Standard | 10.00 / 50.00 USD | 2.00 / 10.00 USD |
| Fast | 20.00 / 100.00 USD | 4.00 / 20.00 USD |
| Ultrafast | 60.00 / 300.00 USD | 12.00 / 60.00 USD |
Cached input on Ultrafast is 6.00 US dollars per million for Astra and 0.60 for GPT-6.1 Sol. Prompts above 272,000 input tokens cost more on both: 120 and 450 for Astra, 24 and 90 for GPT-6.1 Sol.
The practical consequence is that GPT-6.1 Sol Ultrafast, at 12 and 60 US dollars, now costs only modestly more than GPT-6 Astra at Standard speed (10 and 50). For a latency-sensitive agent where Sol's quality is enough, that changes the decision: you can buy speed on the mid tier for about the price of the flagship at normal speed.
OpenAI does not publish a speed multiple for either model on its Ultrafast documentation. When it first announced Ultrafast in August 2026 it described a limited-preview tier for GPT-5.6 Sol that "runs up to 14x faster than Standard processing", but that figure was for a different model. Measure the latency you actually get before paying six times the rate for it.
Why Does Data Residency Now Decide Which Model You Use?
Because the two models now differ on it. OpenAI states that "Ultrafast mode for GPT-6.1 Sol supports US and EU data residency and global processing. GPT-6 Astra Ultrafast supports US data residency and global processing only." Its GPT-6 guidance also now says Fast mode "supports EU data residency for GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna, but not GPT-6 Astra".
For an Indian agency with European clients, or any business that has promised EU-only processing, that settles it: if you need speed and EU residency together, GPT-6.1 Sol is the only GPT-6 option at either faster tier. Astra's faster tiers remain outside an EU residency commitment.
Speed tiers tend to get switched on during a demo, when latency is embarrassing, and then left on. Ultrafast is six times the price on both models, and on Astra it sits outside EU residency. Make it an explicit, documented choice per workflow, never a global default.
What Are the Default Ultrafast Rate Limits?
On 6 October 2026 OpenAI simplified its API usage tiers from five to three: Build, Launch and Grow. Ultrafast has separate rate limits from Standard and Fast, and GPT-6.1 Sol's are considerably higher than Astra's.
| API usage tier | GPT-6.1 Sol Ultrafast (TPM) | GPT-6 Astra Ultrafast (TPM) |
|---|---|---|
| Build | 1,000,000 | 500,000 |
| Launch | 4,000,000 | 1,000,000 |
| Grow | 40,000,000 | 5,000,000 |
OpenAI's rate limits guide sets the tier thresholds by total credit purchases: Build at 5 US dollars, Launch at 100, and Grow at 500, with monthly usage limits of 500, 5,000 and 200,000 US dollars respectively. Organisations working with an OpenAI account team can request higher Ultrafast limits.
In ChatGPT, OpenAI says Ultrafast is available on Pro 500 USD and eligible Enterprise and Edu plans, with plan and workspace requirements documented separately. That ChatGPT-side detail sits outside this API guide.
How Do You Turn Ultrafast On?
Set the service tier in each request. OpenAI's HTTP example is a single extra field; swap the model for gpt-6.1-sol to use the mid tier:
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-astra",
"input": "Explain why the sky is blue in one sentence.",
"service_tier": "ultrafast"
}'
OpenAI "strongly recommend[s]" WebSockets for agentic applications that make many tool calls in quick succession, because "without a persistent connection, network overhead can reduce the latency gains". The speed you pay for can disappear into connection setup if your integration opens a fresh HTTP request for every tool call.
When Is Ultrafast Worth Paying For in 2026?
Only when a person is waiting on a multi-step agent and the wait is costing you something. The October change widens the sensible options, because the mid tier now offers it too.
| Workflow | Ultrafast in 2026? | Better option |
|---|---|---|
| Live sales or support agent a customer is waiting on | Possibly, if measured latency is the problem | Try GPT-6.1 Sol Fast first, then Sol Ultrafast before Astra Ultrafast. |
| Latency-sensitive agent under an EU residency commitment | Only on GPT-6.1 Sol | Astra Ultrafast is US and global only. |
| Overnight reporting and content pipelines | No | Batch or Flex at half the Standard rate. |
| High-volume classification or extraction | No | GPT-6 Luna at 0.10 USD input. |
| An internal demo that feels slow | No | Fix the connection model first; use WebSockets. |
What Are the Common Mistakes With Ultrafast in 2026?
- Defaulting to Astra Ultrafast. GPT-6.1 Sol Ultrafast costs a fifth as much and has higher default limits.
- Assuming the 14x figure applies. That claim was made for GPT-5.6 Sol's limited preview, not either GPT-6 model.
- Leaving it on after a demo. It costs six times Standard on both models.
- Enabling Astra Ultrafast under an EU residency commitment. Only GPT-6.1 Sol supports EU residency at this tier.
- Using HTTP for tool-heavy agents. OpenAI recommends WebSockets.
- Quoting old tier names. Since 6 October 2026 the tiers are Build, Launch and Grow.
Key Takeaways for 2026
- Ultrafast launched for GPT-6 Astra on 29 September 2026 and for GPT-6.1 Sol on 8 October 2026.
- It costs six times Standard on both: 60 and 300 US dollars per million tokens for Astra, 12 and 60 for GPT-6.1 Sol.
- GPT-6.1 Sol Ultrafast supports US and EU data residency; Astra Ultrafast supports US residency and global processing only.
- Default limits run from 1,000,000 to 40,000,000 tokens per minute for Sol, and 500,000 to 5,000,000 for Astra, across the new Build, Launch and Grow tiers.
- OpenAI does not publish a speed multiple for either model. Measure before you pay.
- Use it only where a person is waiting on a multi-step agent, start with the mid tier, and use WebSockets.
Distk helps growth teams across India and internationally decide which workflows actually need faster inference, and which can run on cheaper tiers or overnight batches without anyone noticing. If someone on your team wants to switch on Ultrafast in 2026, that measurement is worth doing first.
Sources
- OpenAI developer changelog, entries for 13 August, 29 September, 6 October and 8 October 2026.
- OpenAI, Ultrafast mode guide.
- OpenAI, GPT-6.1 Sol model page.
- OpenAI API pricing.
- OpenAI, Using GPT-6.
- OpenAI, Rate limits and usage tiers.
- OpenAI, Codex models.