Which AI Models Are Being Compared in September 2026?
Six models from four vendors, all current as of 13 September 2026, covering three tiers. GPT-6 Astra and Claude Fable 5.1 are the frontier tier, both released in September 2026. Gemini 3.8 Flash is Google's generally available workhorse, Gemini 3.1 Pro is its preview flagship with 3.5 Pro announced as coming soon, and Gemini 3.5 Flash-Lite is its volume tier. DeepSeek V4.1 Flash, released 10 September 2026, is DeepSeek's only current-generation model and is priced as a volume tier while benchmarking as a workhorse. Each has its own guide on this site, linked throughout; this page is the one-screen view.
| Model | Vendor | Tier | Status, Sep 2026 | Guide |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI | Frontier | Rolling out; Enterprise off by default | GPT-6 Astra guide |
| Claude Fable 5.1 | Anthropic | Frontier | Generally available; Mythos 5.1 trusted access only | Fable 5.1 guide |
| Gemini 3.1 Pro | Frontier (preview) | Preview; 3.5 Pro coming soon | Gemini 3.1 Pro guide | |
| Gemini 3.8 Flash | Workhorse | Generally available | Gemini 3.8 Flash guide | |
| Gemini 3.5 Flash-Lite | Volume | Generally available | Gemini 3.5 Flash-Lite guide | |
| DeepSeek V4.1 Flash | DeepSeek | Volume price, workhorse benchmarks | Generally available; MIT weights | DeepSeek V4.1 Flash guide |
What Does Each AI Model Cost per Million Tokens in September 2026?
The table below reproduces each vendor's published list rate as of 13 September 2026, in US dollars per million tokens. Where a vendor's model page or launch post does not state a rate, the cell says so rather than borrowing a number from a previous model or a third-party site. DeepSeek rows show peak and off-peak; off-peak is half of peak and applies outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.
| Model | Input (uncached) | Cached input | Output | Notes from the vendor |
|---|---|---|---|---|
| GPT-6 Astra, Standard | 10.00 | Separate rates for cache reads and writes, not enumerated at launch | 50.00 | Fast mode at 2x speed for 2x price (20.00 / 100.00). Zero data retention for eligible API customers. |
| Claude Fable 5.1 | 10.00 | 0.25 (cache read, cut 75 percent) | 50.00 | Anthropic estimates 25 percent lower typical cost than Fable 5, up to 45 percent on agentic work. Text watermark on outputs. |
| Gemini 3.1 Pro | Not published on model page | Not published | Not published on model page | Preview status. Check the developer pricing page. |
| Gemini 3.8 Flash | Not published on model page | Not published | Not published on model page | The 3.7 Flash introductory rate (0.75 / 3.75, expiring 31 December 2026) was model-specific; do not assume it carries over. |
| Gemini 3.5 Flash-Lite | 0.30 | Not stated on model page | 2.50 | Up from 3.1 Flash-Lite (0.25 / 1.50). 350 output tokens per second per Artificial Analysis. |
| DeepSeek V4.1 Flash, peak | 0.30 | 0.006 (cache hit) | 1.20 | Peak is 06:30 to 09:30 and 11:30 to 15:30 IST on weekdays. Concurrency 2,500. |
| DeepSeek V4.1 Flash, off-peak | 0.15 | 0.003 (cache hit) | 0.60 | All other hours and weekends. Exactly half of peak. |
| DeepSeek V4 Pro 0813, peak | 1.32 | 0.044 (cache hit) | 3.96 | Phase-out announced 10 September, continuation announced 11 September. Previous generation. |
| GPT-5.6 Sol | 5.00 | 90 percent discount on cache reads; writes at 1.25x | 30.00 | Previous OpenAI flagship, July 2026, for reference. Terra 2.50 / 15.00, Luna 1.00 / 6.00. |
Three readings of the pricing table
- The frontier has one price. Astra and Fable 5.1 list identically at 10 and 50. The cost difference between them now sits in cache behaviour, token efficiency and effort settings, none of which show on a rate card. Anthropic published a cache-read rate; OpenAI did not enumerate one at launch.
- The volume tier spans two orders of magnitude on cache. Claude's cache read is 0.25. DeepSeek's cache hit is 0.006 at peak. For agentic work that re-reads its context hundreds of times, this line, not the headline input rate, decides the bill.
- Two of Google's three tiers have no published price on their model pages. Gemini 3.8 Flash and 3.1 Pro must be budgeted from the developer pricing page, and 3.1 Pro is in preview with a successor announced.
What Are the Context Windows, Output Limits and Modalities in 2026?
| Model | Context window | Max output | Input modalities | Notable tool features |
|---|---|---|---|---|
| GPT-6 Astra | Not stated; evaluations to 1M | Not stated | Not enumerated in launch post | Computer use; Sites in ChatGPT; cross-window notes in Codex |
| Claude Fable 5.1 | Not stated in launch post | Not stated | Not enumerated in launch post | Effort levels Low, Medium, High; vulnerability discovery permitted |
| Gemini 3.1 Pro | 1M | 64K | Text, image, video, audio, PDF | Function calling, structured output, Search as a tool, code execution |
| Gemini 3.8 Flash | 1M | 64K | Text, image, video, audio, PDF | Function calling, Search as a tool, computer use |
| Gemini 3.5 Flash-Lite | 1M | 64K | Text, image, video, audio, PDF | Function calling, Search as a tool, computer use; thinking levels |
| DeepSeek V4.1 Flash | 1M | 384K | Text, image | JSON output, tool calls, Responses API, Anthropic-format API; reasoning effort 1 to 100 |
Two details stand out. DeepSeek V4.1 Flash's 384K maximum output is six times Google's 64K, which matters for very long single-response deliverables. And Google's model pages are the most complete on modality and limits, while the two frontier launch posts from OpenAI and Anthropic are the least, focusing on benchmarks, safety and price rather than specifications.
Which AI Model Fits Which Workflow in September 2026?
Every vendor's own comparison table this month tells the same story from a different angle: frontier models lead on open-ended agentic work, unaided expert reasoning and consequential computer use; workhorse and volume models tie or lead on defined, multi-step, verifiable workflows at a small fraction of the price. The table below maps common business workflows to the tier and the candidates, using each vendor's published evidence.
| Workflow | Tier | Candidates in Sep 2026 | Evidence (vendor-reported) | Human checkpoint |
|---|---|---|---|---|
| Classification, extraction, translation at volume | Volume | DeepSeek V4.1 Flash, Gemini 3.5 Flash-Lite | DocVQA 95.6 (DeepSeek); OSWorld-Verified 74.0, 350 tok/s (Google) | Sampled audit |
| Content variants under an orchestrator | Volume | DeepSeek V4.1 Flash, Gemini 3.5 Flash-Lite | AutomationBench 54.8 (DeepSeek) | Editorial pass |
| Document-heavy research and audits | Workhorse | Gemini 3.8 Flash, DeepSeek V4.1 Flash at high effort | Glean reports 3x completed long-document tasks vs 3.7 Flash; 1M context on both | Verify citations |
| Agentic coding on established tasks | Workhorse | DeepSeek V4.1 Flash, Gemini 3.8 Flash | DeepSWE 74.2 (DeepSeek, scaffold-dependent); 73.8 (Gemini, per OpenAI) | Code review |
| Reporting assembly from dashboards and exports | Workhorse | Gemini 3.8 Flash, GPT-6 Astra where computer use is needed | Computer use listed on both | Spot-check figures |
| CRM operations, form filling, browser tasks | Frontier | GPT-6 Astra | OSWorld 2.0 72.6 at 40 min/task; 0 percent scope overreach (OpenAI) | Approval gate on writes |
| Novel agentic builds, hardest terminal tasks | Frontier | GPT-6 Astra, Claude Fable 5.1 | Terminal-Bench 4.0: 57.9 (OpenAI), 55.8 (Anthropic) vs 31.2 (DeepSeek), 19.1 (Gemini 3.8 Flash per OpenAI) | Code review |
| Strategy, positioning, final judgement | Frontier | Claude Fable 5.1, GPT-6 Astra | HLE with tools 65.0 (Anthropic); AA Intelligence Index 65.7 vs 61.2 (per OpenAI) | Not an automation candidate |
| Template-faithful decks, docs, spreadsheets | Frontier | GPT-6 Astra | Named strength in OpenAI's launch post | Claims and numbers review |
| Regulated personal data | Any tier with a published policy | Claude Fable 5.1 (EFS, ZDR), GPT-6 Astra (ZDR), self-hosted DeepSeek | DeepSeek API publishes no retention policy | Procurement sign-off |
What Did Each Vendor Not Publish in September 2026?
The gaps are as informative as the numbers, and a comparison page that hides them is selling something. Each item below is a field a buyer would reasonably want and the vendor's own material does not supply.
- OpenAI, GPT-6 Astra: cache read and write rates, context window, output limit, what Astra Pro adds. Alignment regression on reasoning monitorability disclosed.
- Anthropic, Claude Fable 5.1: context window and output limit in the launch post; whether editing removes the text watermark. Approval-bypass limitation disclosed.
- Google, Gemini 3.8 Flash: any price; scores for three of the four named benchmarks. Only HLE-Verified at 54.9 has a number.
- Google, Gemini 3.1 Pro: any price; a date for 3.5 Pro. Comparison table uses superseded competitor models.
- Google, Gemini 3.5 Flash-Lite: cached input rate. Otherwise the most complete page in the set.
- DeepSeek, V4.1 Flash: data retention, residency or zero-data-retention terms; minimum self-host hardware; a consistent V4 Pro retirement date. Comparison table uses superseded competitor models.
How Should a Business Route Across These Models in 2026?
In three tiers, chosen per workflow, behind an abstraction that references a role rather than a model name, validated by a fixed evaluation set and reviewed quarterly. This is the same recommendation we have made in every model guide this year, and September 2026 is the month that proved why: three vendors shipped new models in one month, one vendor reversed a retirement within 24 hours, and Google's flagship tier is a generation behind its workhorse. A stack that hardcodes a model name is out of date on arrival.
| Tier | Role in the stack | September 2026 default | Alternative | Budget expectation |
|---|---|---|---|---|
| Volume | The worker: high-count, defined, verifiable tasks | DeepSeek V4.1 Flash, scheduled off-peak where possible | Gemini 3.5 Flash-Lite where data policy or Google ecosystem favours it | Cents to low dollars per million tokens |
| Workhorse | Agents, long documents, reporting, builds | Gemini 3.8 Flash | DeepSeek V4.1 Flash at high effort; price of 3.8 Flash to be confirmed | Low single-digit dollars per million, to be confirmed for Gemini |
| Frontier | Judgement, novel agentic work, consequential computer use | Claude Fable 5.1 or GPT-6 Astra by task type | Gemini 3.1 Pro for creative prototypes and research agents, by role, until 3.5 Pro | 10 and 50 dollars per million; cache behaviour decides the real number |
Four operating rules for 2026
- Keep a standing evaluation set of twenty to fifty real tasks per workflow with known good outputs. A new model is an afternoon's test, not a quarter's project.
- Route by role, not by name. The volume model, the workhorse, the frontier. A release becomes a configuration change.
- Measure cost per passing task, not cost per token. Effort settings, retries and cache behaviour all move the real number away from the rate card.
- Answer the data question before the API key. Where does personal data go, on what basis, for how long. Two vendors published an answer this month; one did not.
Where Are the Detailed Guides for Each Model in 2026?
Each row in this comparison has a full guide with the vendor's complete published figures, the honest caveats and the workflow fit in detail.
- GPT-6 Astra in 2026: What OpenAI's Computer-Use Model Means for Business Teams
- Claude Fable 5.1 and Mythos 5.1 in 2026: What Changed for Business Teams
- Gemini 3.8 Flash in 2026: Google's Agent Workhorse Explained
- Gemini 3.5 Flash-Lite in 2026: The High-Volume Tier Explained
- Gemini 3.1 Pro in 2026: The Reasoning Tier, Still in Preview
- What Is DeepSeek V4.1 Flash in 2026? The New Architecture Explained
- DeepSeek V4.1 Flash Pricing in 2026: Rate Card, Time Zones and Real Cost
- DeepSeek V4.1 Flash Benchmarks in 2026: Where It Leads and Trails
- Migrating From DeepSeek V4 Pro to V4.1 Flash in 2026
- DeepSeek V4.1 Flash for Marketing Teams in 2026
- DeepSeek V4.1 Flash Open Weights in 2026: What Self-Hosting Involves
- DeepSeek V4 Pro 0813 in 2026: The August Rate Card Change
- What Is GPT-5.6 in 2026? Sol, Terra and Luna Explained
- Gemini 3.7 Flash in 2026: The Introductory Price and the 2027 Cliff
What Are the Common Mistakes in Comparing AI Models in 2026?
- Comparing on headline input price. Cache behaviour and output rates decide agentic and content bills. Output is five times input at the frontier and four times at DeepSeek.
- Filling unpublished cells with guesses. Two Gemini tiers and OpenAI's cache rates are not on the vendor pages. Say so.
- Reading any vendor's table as neutral. Every comparison this month was run by the vendor, and several use competitor models superseded the same month.
- Choosing one model for everything. The lead changes hands by task type on every table published in September 2026.
- Ignoring status. Preview, rolling out, off by default for Enterprise, retirement announced then paused. Status is a planning input.
- Skipping the data question. Cheapest is not a procurement answer.
Key Takeaways for September 2026
- Frontier price has converged at 10 and 50 US dollars per million tokens for GPT-6 Astra and Claude Fable 5.1. Cache behaviour is the differentiator.
- Volume price spans from Gemini 3.5 Flash-Lite at 0.30 and 2.50 to DeepSeek V4.1 Flash at 0.30 and 1.20 peak, 0.15 and 0.60 off-peak, with cache hits at 0.006.
- Gemini 3.8 Flash and 3.1 Pro publish no price on their model pages; budget from the developer pricing page.
- Fit follows tier on every vendor's table: frontier for judgement and novel agentic work, workhorse and volume for defined multi-step workflows.
- DeepSeek publishes no data-retention policy; Anthropic and OpenAI published zero-data-retention positions.
- Route in three tiers by role, keep an evaluation set, measure cost per passing task, review quarterly.
Distk builds exactly this routed stack for growth teams across India and internationally: the evaluation set, the tier assignments, the effort settings, the off-peak scheduling and the data-handling answer for procurement. If September 2026 has left your model choice out of date, that rebuild is where we start.