AI Model Guide

AI Model Pricing and Fit Comparison, September 2026

In September 2026, OpenAI, Anthropic and DeepSeek all shipped new models and Google's three-tier Gemini lineup sat in between. This page puts every published rate card, context window and fit signal on one screen, flags what the vendors did not publish, and ends with a routing recommendation a business team can act on.

Distk Editorial Sep 2026 14 min read

As of 13 September 2026 the frontier tier has converged: GPT-6 Astra and Claude Fable 5.1 both list at 10 US dollars per million input tokens and 50 per million output, with Anthropic's cache reads at 0.25 and OpenAI's cache rates separate and unenumerated. The workhorse and volume tiers diverge widely: Gemini 3.5 Flash-Lite lists at 0.30 and 2.50, DeepSeek V4.1 Flash at 0.30 and 1.20 at peak and half that off-peak with cache hits at 0.006, and Gemini 3.8 Flash and Gemini 3.1 Pro publish no price on their model pages. Fit follows a pattern across every vendor's own tables: frontier models lead on open-ended agentic work, unaided expert reasoning and computer use with consequences; workhorse and volume models tie or lead on defined multi-step workflows at a fraction of the cost. The routing that follows is three tiers, chosen per workflow, with a fixed evaluation set and a quarterly review.

Which AI Models Are Being Compared in September 2026?

Six models from four vendors, all current as of 13 September 2026, covering three tiers. GPT-6 Astra and Claude Fable 5.1 are the frontier tier, both released in September 2026. Gemini 3.8 Flash is Google's generally available workhorse, Gemini 3.1 Pro is its preview flagship with 3.5 Pro announced as coming soon, and Gemini 3.5 Flash-Lite is its volume tier. DeepSeek V4.1 Flash, released 10 September 2026, is DeepSeek's only current-generation model and is priced as a volume tier while benchmarking as a workhorse. Each has its own guide on this site, linked throughout; this page is the one-screen view.

ModelVendorTierStatus, Sep 2026Guide
GPT-6 AstraOpenAIFrontierRolling out; Enterprise off by defaultGPT-6 Astra guide
Claude Fable 5.1AnthropicFrontierGenerally available; Mythos 5.1 trusted access onlyFable 5.1 guide
Gemini 3.1 ProGoogleFrontier (preview)Preview; 3.5 Pro coming soonGemini 3.1 Pro guide
Gemini 3.8 FlashGoogleWorkhorseGenerally availableGemini 3.8 Flash guide
Gemini 3.5 Flash-LiteGoogleVolumeGenerally availableGemini 3.5 Flash-Lite guide
DeepSeek V4.1 FlashDeepSeekVolume price, workhorse benchmarksGenerally available; MIT weightsDeepSeek V4.1 Flash guide

What Does Each AI Model Cost per Million Tokens in September 2026?

The table below reproduces each vendor's published list rate as of 13 September 2026, in US dollars per million tokens. Where a vendor's model page or launch post does not state a rate, the cell says so rather than borrowing a number from a previous model or a third-party site. DeepSeek rows show peak and off-peak; off-peak is half of peak and applies outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

ModelInput (uncached)Cached inputOutputNotes from the vendor
GPT-6 Astra, Standard10.00Separate rates for cache reads and writes, not enumerated at launch50.00Fast mode at 2x speed for 2x price (20.00 / 100.00). Zero data retention for eligible API customers.
Claude Fable 5.110.000.25 (cache read, cut 75 percent)50.00Anthropic estimates 25 percent lower typical cost than Fable 5, up to 45 percent on agentic work. Text watermark on outputs.
Gemini 3.1 ProNot published on model pageNot publishedNot published on model pagePreview status. Check the developer pricing page.
Gemini 3.8 FlashNot published on model pageNot publishedNot published on model pageThe 3.7 Flash introductory rate (0.75 / 3.75, expiring 31 December 2026) was model-specific; do not assume it carries over.
Gemini 3.5 Flash-Lite0.30Not stated on model page2.50Up from 3.1 Flash-Lite (0.25 / 1.50). 350 output tokens per second per Artificial Analysis.
DeepSeek V4.1 Flash, peak0.300.006 (cache hit)1.20Peak is 06:30 to 09:30 and 11:30 to 15:30 IST on weekdays. Concurrency 2,500.
DeepSeek V4.1 Flash, off-peak0.150.003 (cache hit)0.60All other hours and weekends. Exactly half of peak.
DeepSeek V4 Pro 0813, peak1.320.044 (cache hit)3.96Phase-out announced 10 September, continuation announced 11 September. Previous generation.
GPT-5.6 Sol5.0090 percent discount on cache reads; writes at 1.25x30.00Previous OpenAI flagship, July 2026, for reference. Terra 2.50 / 15.00, Luna 1.00 / 6.00.

Three readings of the pricing table

What Are the Context Windows, Output Limits and Modalities in 2026?

ModelContext windowMax outputInput modalitiesNotable tool features
GPT-6 AstraNot stated; evaluations to 1MNot statedNot enumerated in launch postComputer use; Sites in ChatGPT; cross-window notes in Codex
Claude Fable 5.1Not stated in launch postNot statedNot enumerated in launch postEffort levels Low, Medium, High; vulnerability discovery permitted
Gemini 3.1 Pro1M64KText, image, video, audio, PDFFunction calling, structured output, Search as a tool, code execution
Gemini 3.8 Flash1M64KText, image, video, audio, PDFFunction calling, Search as a tool, computer use
Gemini 3.5 Flash-Lite1M64KText, image, video, audio, PDFFunction calling, Search as a tool, computer use; thinking levels
DeepSeek V4.1 Flash1M384KText, imageJSON output, tool calls, Responses API, Anthropic-format API; reasoning effort 1 to 100

Two details stand out. DeepSeek V4.1 Flash's 384K maximum output is six times Google's 64K, which matters for very long single-response deliverables. And Google's model pages are the most complete on modality and limits, while the two frontier launch posts from OpenAI and Anthropic are the least, focusing on benchmarks, safety and price rather than specifications.

Which AI Model Fits Which Workflow in September 2026?

Every vendor's own comparison table this month tells the same story from a different angle: frontier models lead on open-ended agentic work, unaided expert reasoning and consequential computer use; workhorse and volume models tie or lead on defined, multi-step, verifiable workflows at a small fraction of the price. The table below maps common business workflows to the tier and the candidates, using each vendor's published evidence.

WorkflowTierCandidates in Sep 2026Evidence (vendor-reported)Human checkpoint
Classification, extraction, translation at volumeVolumeDeepSeek V4.1 Flash, Gemini 3.5 Flash-LiteDocVQA 95.6 (DeepSeek); OSWorld-Verified 74.0, 350 tok/s (Google)Sampled audit
Content variants under an orchestratorVolumeDeepSeek V4.1 Flash, Gemini 3.5 Flash-LiteAutomationBench 54.8 (DeepSeek)Editorial pass
Document-heavy research and auditsWorkhorseGemini 3.8 Flash, DeepSeek V4.1 Flash at high effortGlean reports 3x completed long-document tasks vs 3.7 Flash; 1M context on bothVerify citations
Agentic coding on established tasksWorkhorseDeepSeek V4.1 Flash, Gemini 3.8 FlashDeepSWE 74.2 (DeepSeek, scaffold-dependent); 73.8 (Gemini, per OpenAI)Code review
Reporting assembly from dashboards and exportsWorkhorseGemini 3.8 Flash, GPT-6 Astra where computer use is neededComputer use listed on bothSpot-check figures
CRM operations, form filling, browser tasksFrontierGPT-6 AstraOSWorld 2.0 72.6 at 40 min/task; 0 percent scope overreach (OpenAI)Approval gate on writes
Novel agentic builds, hardest terminal tasksFrontierGPT-6 Astra, Claude Fable 5.1Terminal-Bench 4.0: 57.9 (OpenAI), 55.8 (Anthropic) vs 31.2 (DeepSeek), 19.1 (Gemini 3.8 Flash per OpenAI)Code review
Strategy, positioning, final judgementFrontierClaude Fable 5.1, GPT-6 AstraHLE with tools 65.0 (Anthropic); AA Intelligence Index 65.7 vs 61.2 (per OpenAI)Not an automation candidate
Template-faithful decks, docs, spreadsheetsFrontierGPT-6 AstraNamed strength in OpenAI's launch postClaims and numbers review
Regulated personal dataAny tier with a published policyClaude Fable 5.1 (EFS, ZDR), GPT-6 Astra (ZDR), self-hosted DeepSeekDeepSeek API publishes no retention policyProcurement sign-off

What Did Each Vendor Not Publish in September 2026?

The gaps are as informative as the numbers, and a comparison page that hides them is selling something. Each item below is a field a buyer would reasonably want and the vendor's own material does not supply.

How Should a Business Route Across These Models in 2026?

In three tiers, chosen per workflow, behind an abstraction that references a role rather than a model name, validated by a fixed evaluation set and reviewed quarterly. This is the same recommendation we have made in every model guide this year, and September 2026 is the month that proved why: three vendors shipped new models in one month, one vendor reversed a retirement within 24 hours, and Google's flagship tier is a generation behind its workhorse. A stack that hardcodes a model name is out of date on arrival.

TierRole in the stackSeptember 2026 defaultAlternativeBudget expectation
VolumeThe worker: high-count, defined, verifiable tasksDeepSeek V4.1 Flash, scheduled off-peak where possibleGemini 3.5 Flash-Lite where data policy or Google ecosystem favours itCents to low dollars per million tokens
WorkhorseAgents, long documents, reporting, buildsGemini 3.8 FlashDeepSeek V4.1 Flash at high effort; price of 3.8 Flash to be confirmedLow single-digit dollars per million, to be confirmed for Gemini
FrontierJudgement, novel agentic work, consequential computer useClaude Fable 5.1 or GPT-6 Astra by task typeGemini 3.1 Pro for creative prototypes and research agents, by role, until 3.5 Pro10 and 50 dollars per million; cache behaviour decides the real number

Four operating rules for 2026

  1. Keep a standing evaluation set of twenty to fifty real tasks per workflow with known good outputs. A new model is an afternoon's test, not a quarter's project.
  2. Route by role, not by name. The volume model, the workhorse, the frontier. A release becomes a configuration change.
  3. Measure cost per passing task, not cost per token. Effort settings, retries and cache behaviour all move the real number away from the rate card.
  4. Answer the data question before the API key. Where does personal data go, on what basis, for how long. Two vendors published an answer this month; one did not.

Where Are the Detailed Guides for Each Model in 2026?

Each row in this comparison has a full guide with the vendor's complete published figures, the honest caveats and the workflow fit in detail.

What Are the Common Mistakes in Comparing AI Models in 2026?

Key Takeaways for September 2026

Distk builds exactly this routed stack for growth teams across India and internationally: the evaluation set, the tier assignments, the effort settings, the off-peak scheduling and the data-handling answer for procurement. If September 2026 has left your model choice out of date, that rebuild is where we start.

AI Model Pricing in September 2026: FAQs

What is the cheapest capable AI model in September 2026?

On published list rates, DeepSeek V4.1 Flash: 0.30 US dollars per million cache-miss input tokens and 1.20 per million output at peak, half that off-peak, with cache hits at 0.006. Gemini 3.5 Flash-Lite is the closest published alternative at 0.30 and 2.50.

How much do GPT-6 Astra and Claude Fable 5.1 cost?

Both list at 10 US dollars per million input tokens and 50 per million output. Claude's cache reads are 0.25; OpenAI's cache rates are separate and not enumerated in the launch post. GPT-6 Astra Fast mode is 20 and 100.

How much does Gemini 3.8 Flash cost?

Google's 3.8 Flash model page does not publish a price, and neither does the 3.1 Pro page. Gemini 3.5 Flash-Lite lists at 0.30 and 2.50. Budget the other two from Google's developer pricing page rather than from the 3.7 Flash introductory rate.

Which AI model is best for marketing teams in 2026?

No single one. Volume work fits DeepSeek V4.1 Flash or Gemini 3.5 Flash-Lite; document-heavy agents and reporting fit Gemini 3.8 Flash; strategy, judgement and consequential computer use fit Claude Fable 5.1 or GPT-6 Astra. Route by workflow.

Which vendors publish a data-retention policy?

Anthropic (Enterprise Frontier Safeguards and zero data retention for eligible customers) and OpenAI (zero data retention for eligible API customers) do. DeepSeek's V4.1 Flash pages publish no retention or residency terms. Google's model pages do not address it.

Are the benchmark comparisons between these models reliable?

Every comparison published in September 2026 was run by the vendor, and both DeepSeek's and Google's Pro tables compare against competitor models that were superseded the same month. Treat them as directional and test on your own evaluation set.

Route by role, review quarterly, measure cost per passing task

Distk builds the three-tier routed stack, the evaluation set and the data-handling answer that turn September 2026's rate cards into a decision you can defend to finance and procurement. When the next model ships, it is an afternoon's test.

Start the conversation →