What Is GPT-6.1 Sol in 2026?
GPT-6.1 Sol is an upgrade to GPT-6 Sol announced by OpenAI on 29 September 2026. OpenAI's own framing is the headline: near-Astra intelligence for a fifth of the price. It nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one fifth of Astra's standard input and output token prices. It is available to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, and through the API as gpt-6.1-sol.
One availability detail catches people out. GPT-6.1 Sol is not yet available in Chat. It lives in ChatGPT Work and Codex, which is consistent with the Sol line being positioned as the agentic workhorse rather than the conversational model.
How Much Does GPT-6.1 Sol Cost in 2026?
Standard API pricing is 2 US dollars per million input tokens, 0.10 US dollars per million cached input tokens and 10 US dollars per million output tokens. The cached input rate is the number to notice: OpenAI states it is 95 percent below standard input pricing and 50 percent below GPT-6 Sol's cached input rate, which is aimed squarely at agents that reuse context across requests.
| Price per 1M tokens | GPT-6.1 Sol | GPT-6 Astra (Standard) | Claude Sonnet 5.5 |
|---|---|---|---|
| Input | 2.00 USD | 10.00 USD | 2.00 USD |
| Cached input | 0.10 USD | Separate rates, not enumerated at launch | 0.20 USD (cache read) |
| Output | 10.00 USD | 50.00 USD | 10.00 USD |
That puts GPT-6.1 Sol at exactly the same headline rate as Claude Sonnet 5.5, with a lower cached input rate. The mid tier across both vendors has now converged on 2 and 10 US dollars, and the competition has moved to cached context and tokens consumed per task. Our model pricing comparison holds the full cross-vendor table.
OpenAI also says GPT-6.1 Sol Ultrafast is coming in the following days, with up to 8x faster token generation compared with its standard speed in Codex.
What Do the GPT-6.1 Sol Benchmarks Actually Say in 2026?
OpenAI published cost-per-task comparisons rather than a single score table, which is more useful and harder to quote out of context. The pattern across five evaluations is consistent: close to Astra on capability, a fraction of the cost, and in two cases ahead of Claude Opus 5.5.
| Evaluation | What OpenAI reports for GPT-6.1 Sol in 2026 | What it measures |
|---|---|---|
| DeepSWE v1.1 | Matches GPT-6 Astra at roughly one fifth of the cost, and beats GPT-6 Sol's best score by 6.4 points at a lower reasoning effort and cost | Complex software engineering in real codebases |
| GDP.pdf | Scores higher than Opus 5.5 with fallbacks at less than half the cost per task, and approaches Astra at roughly one fifth the cost per task | Professional questions over complex PDFs with tables, charts and fine print |
| AutomationBench | 2.2 points above Opus 5.5 at medium reasoning effort at roughly a third of the cost, and 4.8 points above GPT-6 Sol at the same setting | Multi-step business workflows across 47 tools |
| OSWorld 2.0 (offline) | Seven points above GPT-6 Sol at maximum effort at less than half the cost, and within 2.1 points of Astra at roughly one seventh the cost per task | Long-horizon computer use |
| Terminal-Bench Science 0.1 | More than doubles GPT-6 Sol at maximum effort at less than half the cost per task; 5.47 USD per task against 23.21 for Opus 5.5 and 23.80 for Astra | Scientific workflows, data analysis, simulation, theorem proving |
| Factuality | At low reasoning effort, responses containing a factual error fall from 11.4 percent to 7.7 percent, about a 32 percent reduction | Share of answers with at least one factual error on deliberately difficult prompts |
The Terminal-Bench Science row is the clearest illustration of the pitch: 5.47 US dollars per task against roughly 23 for both Opus 5.5 and Astra, which OpenAI describes as over 75 percent lower cost than either. It also states plainly that Astra still achieves the highest score among models tested at 68.1 percent and should be used for the most difficult scientific research tasks.
The caveats OpenAI published alongside the numbers
- AutomationBench comparison note: OpenAI states the Claude Fable 5.1 datapoint understates its actual cost, because it omits the cost of fallbacks, which occurred on roughly 40 percent of tasks.
- The Opus 5.5 comparison includes fallbacks on GDP.pdf, which is the honest way to price a model whose safeguards route work elsewhere.
- Factuality prompts are adversarial: the evaluation uses de-identified ChatGPT conversations where users had flagged an earlier model's factual error. OpenAI says these error-inducing prompts are not representative of typical usage.
- Effort settings vary by row, so a figure at maximum effort and one at medium effort are not comparable without reading which is which.
- Every figure is OpenAI's own, including the competitor numbers.
Why Does the Cached Input Price Matter So Much in 2026?
Because agentic work re-reads the same context on every step. A task that runs forty tool calls sends its system prompt, its tool definitions and its accumulated notes forty times. At 0.10 US dollars per million cached input tokens, that repetition becomes close to free, which is precisely why OpenAI frames it as giving developers more room to build and run capable agents that reuse context across requests.
This is now the main axis of competition in the mid tier. Anthropic cut Fable 5.1's cache reads to 0.25 and set both Opus 5.5 and Sonnet 5.5 at 0.20. OpenAI has gone to 0.10. If your workload is agentic, the cached rate will move your bill more than the headline input rate, and it is the number to put in a budget model first.
Two vendors now offer a mid-tier model at 2 and 10 US dollars per million tokens with sub-0.25 cached input. At that price the deciding factors are no longer rate cards but tokens consumed per task, how your effort setting is configured, and whether safeguards route some of your work to a fallback model you also pay for. All three are measurable on your own workload and none of them are visible on a pricing page.
What Does the Alignment Work Mean for Delegation in 2026?
OpenAI reports substantial improvements over GPT-6 Sol on its alignment evaluations, bringing GPT-6.1 Sol closer to Astra. Three specifics matter if you plan to run it unattended: it is more transparent about its limitations, more reliable at respecting user intent and safety constraints, and it shows lower failure rates on transparency about broken search tools, respecting explicit restrictions, and avoiding unauthorised outcomes during agentic tasks.
The broken-search-tool evaluation is the most practical one published, because it tests whether an agent admits a tool has failed instead of guessing. GPT-6.1 Sol fails to disclose the problem in 2.1 percent of cases, against 4.9 percent for GPT-6 Sol, 1.5 percent for GPT-6 Astra and 28.7 percent for GPT-6 Luna. That Luna figure is worth remembering if you were considering the cheapest tier for anything where a silent failure would be costly. OpenAI also reports no attempts to bypass an automated safety reviewer, matching both Astra and GPT-6 Sol, and notes these evaluations deliberately test challenging situations rather than measuring typical use.
How Should Marketing and Growth Teams Use GPT-6.1 Sol in 2026?
Make it the default for agentic and document work, and keep Astra for the hardest reasoning. The three evaluations closest to ordinary business work, AutomationBench, GDP.pdf and OSWorld, are exactly where OpenAI reports it beating or nearly matching much more expensive models. That is an unusual combination and it changes which workflows are affordable to run repeatedly rather than occasionally.
| Workflow | Model choice in 2026 | Human checkpoint |
|---|---|---|
| Multi-step business workflow automation | GPT-6.1 Sol. Above Opus 5.5 on AutomationBench at a third of the cost. | Sampled audit; approval on anything outbound. |
| Extracting answers from contracts, invoices and reports | GPT-6.1 Sol. This is the GDP.pdf use case. | Verify figures against the source document. |
| Long-horizon computer use and browser tasks | GPT-6.1 Sol at high effort, within 2.1 points of Astra at a seventh of the cost. | No unattended writes to live systems. |
| Agents that reuse a large standing context | GPT-6.1 Sol. The 0.10 cached rate is the whole argument. | Instrument token usage per workflow. |
| Landing pages, tracking fixes, internal tooling | GPT-6.1 Sol. Matches Astra on DeepSWE at a fifth of the cost. | Code review and staged deploy. |
| Hardest reasoning, research and judgement calls | GPT-6 Astra. OpenAI says so for the most difficult tasks. | Not an automation candidate. |
| Conversational use in ChatGPT Chat | Not available. GPT-6.1 Sol is in Work and Codex only. | Use a Chat-available model. |
What Are the Common Mistakes to Avoid in 2026?
- Looking for it in Chat. GPT-6.1 Sol is not yet available there. It is in ChatGPT Work and Codex, and in the API.
- Comparing figures across effort settings. Some results are at maximum effort, some at medium. The cost claims move with the setting.
- Quoting the factuality improvement as a general accuracy claim. The prompts were selected because a previous model got them wrong.
- Ignoring fallback costs when comparing vendors. OpenAI notes Fable 5.1's AutomationBench cost omits fallbacks that occurred on around 40 percent of tasks, which is exactly the kind of hidden cost to check on your own stack.
- Budgeting on the standard input rate. For agentic work the 0.10 cached input rate will dominate. Instrument which of your tokens are cached.
- Assuming Ultrafast is available now. OpenAI says it is coming in the following days, with up to 8x faster generation in Codex.
- Reaching for the cheapest tier for anything sensitive. GPT-6 Luna failed to disclose a broken search tool in 28.7 percent of cases on OpenAI's own evaluation.
Key Takeaways for 2026
GPT-6.1 Sol is the clearest sign that 2026's real competition is cost per completed task rather than leaderboard position, and that the mid tier is now where most business work belongs.
- Announced 29 September 2026. Available to Plus, Pro, Business, Enterprise and Edu in ChatGPT Work and Codex, and via the API as
gpt-6.1-sol. Not yet in Chat. - Priced at 2 and 10 US dollars per million tokens with cached input at 0.10, which OpenAI says is 95 percent below standard input and half GPT-6 Sol's cached rate.
- Matches Astra on DeepSWE at about a fifth of the cost, and comes within 2.1 points of Astra on OSWorld at roughly a seventh.
- Beats Opus 5.5 on AutomationBench by 2.2 points at a third of the cost, and on GDP.pdf at less than half the cost per task.
- On Terminal-Bench Science it costs 5.47 USD per task against roughly 23 for Opus 5.5 and Astra, while Astra still holds the top score at 68.1 percent.
- Factual error rate at low effort falls from 11.4 to 7.7 percent, on deliberately difficult prompts.
- Alignment improves over GPT-6 Sol, with a 2.1 percent failure rate on disclosing a broken search tool against 28.7 percent for GPT-6 Luna.
- GPT-6.1 Sol Ultrafast, up to 8x faster generation in Codex, is said to follow within days.
Distk helps growth teams across India and internationally decide which workflows belong on a mid tier and which genuinely need a flagship, measure cost per completed task rather than per token, and set the effort level and caching strategy that make an agentic workload affordable to run every day. If GPT-6.1 Sol is on your 2026 shortlist, that measurement is where we start.
Sources
- OpenAI, Introducing GPT-6.1 Sol, 29 September 2026. Every price, benchmark claim, cost-per-task figure and alignment result in this guide comes from that announcement and is OpenAI's own reporting, including the competitor figures.