AI Models

DeepSeek V4 Pro 0813 in 2026: The Rate Card Expires This Week

If you chose this model because it was the cheapest capable option, the number you chose it on stops applying at 16:00 UTC on 16 August 2026. Peak and off-peak billing replaces the flat rate, and one line of the card rises about twelve times. This is a two day problem, and it is worth an hour of arithmetic.

Distk Editorial 14 August 2026 11 min read

DeepSeek V4 Pro 0813 reached general availability on 13 August 2026 and is among the cheapest capable models on the market in 2026. On DeepSeek's own pricing page, that changes at 16:00 UTC on 16 August 2026, when the flat rate is replaced by peak and off-peak billing with off-peak set at half of peak. Cache-hit input rises roughly 12x at peak, cache-miss input roughly 3x and output roughly 4.6x. Two things follow. First, anyone who selected this model on price needs to re-run the numbers this week. Second, the headline benchmark gains circulating in 2026 are DeepSeek measuring itself against its own preview build, not against the alternative you are choosing between. The independent index that does exist puts V4 Pro one point above DeepSeek's own cheaper Flash model.

Why Does the DeepSeek Rate Card Change This Week in 2026?

Because DeepSeek is retiring flat per-token pricing in favour of time-of-day billing. Its API pricing page states that the new prices take effect at 16:00 UTC on 16 August 2026, that pricing moves to peak and off-peak billing, and that off-peak rates are half the peak rates. Peak hours in 2026 are 01:00 to 04:00 and 06:00 to 10:00 UTC. Every other hour is off-peak.

Much of the coverage in 2026 reports this as a 17 August change. That is not a contradiction. DeepSeek is based in China, and 16:00 UTC on 16 August is midnight on 17 August in Beijing. The version of the date that matters is the one in your own timezone, because that is when your invoice starts behaving differently.

The deadline in your timezone

The cutover is a single global instant, so it lands on a different clock face depending on where your finance team sits. For teams in India it arrives on the evening of 16 August 2026, not the 17th. Working backwards from the published UTC timestamp is the only reliable way to know how long you actually have.

LocationWhen the new rate card starts
UTC16:00, 16 August 2026
India (IST)21:30, 16 August 2026
Beijing (CST)00:00, 17 August 2026
London (BST)17:00, 16 August 2026
New York (EDT)12:00, 16 August 2026

What the new rate card says

The new card splits every existing line into a peak rate and an off-peak rate, with off-peak fixed at exactly half of peak. The comparison below sets DeepSeek's published pre-change rates against the published peak and off-peak rates for V4 Pro. The multiplier column is straightforward division of the new figure by the old one, and it is the number most teams have not yet worked out.

Rate line (per 1M tokens)Before 16 Aug 2026Off-peakPeakPeak multiplier
Input, cache hit0.003625 USD0.022 USD0.044 USDabout 12.1x
Input, cache miss0.435 USD0.66 USD1.32 USDabout 3.0x
Output0.87 USD1.98 USD3.96 USDabout 4.6x

This is where the widely repeated 12x headline comes from, and it deserves precision. Cache-hit input rising from 0.003625 to 0.044 US dollars is roughly 12.1 times, and it is the largest single multiplier on the card. It is not the multiplier on your bill unless your traffic is almost entirely cache hits, which is unusual. Quoting 12x as the overall increase overstates what most teams will pay.

Dated note

This section was written on 14 August 2026 against DeepSeek's published pricing page as it stood on that date. Rate cards in 2026 have moved repeatedly across every major vendor, and a scheduled change can be delayed, revised or superseded. Check the live pricing page before you commit a budget to any figure here.

What Does the Rate Change Do to a Real Workload in 2026?

It roughly doubles to roughly quadruples the monthly bill, depending on cache behaviour and when your jobs run. The illustration below uses DeepSeek's published rates and a set of volumes we invented for the purpose. The volumes are not a forecast, a survey result or a client figure. They exist only to show how the multipliers combine, and you should substitute your own usage before drawing any conclusion.

Assume a mid-sized content and support workload of 40 million cache-miss input tokens, 60 million cache-hit input tokens and 12 million output tokens in a month. That shape favours a heavy system prompt reused across many calls, which is typical of marketing automation in 2026.

Scenario (illustrative volumes)Monthly costVersus current
Current rate card, before 16 Aug 202628.06 USDbaseline
All traffic in off-peak hours51.48 USDabout 1.8x
Half peak, half off-peak77.22 USDabout 2.8x
All traffic in peak hours102.96 USDabout 3.7x

The interesting row for Indian teams is the middle one, and it is not a coincidence. Peak hours of 01:00 to 04:00 and 06:00 to 10:00 UTC translate to 06:30 to 09:30 and 11:30 to 15:30 IST. A conventional 09:00 to 18:00 Indian working day therefore sits inside a peak window for four and a half of its nine hours. Interactive workloads driven by staff at their desks in 2026 will land close to a fifty-fifty split by default.

That points at the cheapest available mitigation. Batch work with no human waiting on it, such as overnight content generation, bulk classification and reporting, can be scheduled into off-peak hours at half the peak rate without changing application logic. Interactive traffic cannot be moved, so the saving is partial.

What Is DeepSeek V4 Pro 0813 in 2026?

DeepSeek V4 Pro 0813 is the general availability build of DeepSeek's flagship mixture-of-experts model, rolled out on 13 August 2026 across the DeepSeek app, web and API. It supersedes the earlier preview version, adds substantially stronger agent behaviour according to the vendor, introduces selectable thinking effort levels and ships native support for the OpenAI Responses API format.

The published specification is unusually generous at the price. A one million token context window and a 384,000 token output ceiling put it in the same structural class as far more expensive models in 2026, which is precisely why cost-led buyers picked it up so quickly.

SpecificationDocumented value in 2026
API model identifierdeepseek-v4-pro (version DeepSeek-V4-Pro-0813)
Context window1,048,576 tokens
Maximum output384,000 tokens
Total parameters1.6 trillion, mixture-of-experts
Active parameters per tokenroughly 49 billion
Pretraining corpusmore than 32 trillion tokens
Thinking effort levelslow, high, max
Other capabilitiesthinking mode, JSON output, tool calling
Licence for released weightsMIT

One detail matters for anyone comparing tiers. DeepSeek also sells V4 Flash, whose 0731 build carries the same one million token context and 384,000 token output ceiling at a fraction of the price. The tier decision in 2026 is therefore not about context length, which is the assumption many teams start from.

Why Should You Not Read the Headline Benchmarks as a Buying Signal in 2026?

Because almost every headline number is DeepSeek comparing the new build against its own previous build. That is a legitimate way to report progress on a point release, and it is what most vendors do in 2026. It answers the question of whether the engineering team improved something. It does not answer the question a buyer is actually asking, which is whether this model beats the specific alternative under consideration.

Put plainly: if last year you ran a hundred metres in twenty seconds and this year you run it in twelve, that is a real and impressive improvement. It tells you nothing about whether you would win the race you have entered. The self-comparison is a measure of the vendor's progress, not of your options.

The figures below are the vendor-reported preview-to-GA gains circulating in 2026. DeepSeek published them as an image in its announcement rather than as text, so these come from secondary coverage transcribing that image rather than from a table we could read directly. Treat them as vendor-reported and directional.

BenchmarkV4 Pro PreviewV4 Pro 0813Change
DeepSWE12.862.7+49.9
DSBench-Hard31.167.2+36.1
CyberGym52.783.3+30.6
AutomationBench (Public)12.831.8+19.0
Terminal Bench 2.172.187.9+15.8
Toolathlon-Verifiednot published74.1no baseline given
DSBench-FullStacknot published71.1no baseline given
NL2Reponot published61.5no baseline given

Two of those rows deserve a flag. DSBench-Hard and DSBench-FullStack are described in the surrounding 2026 coverage as DeepSeek's own internal test sets rather than public third-party benchmarks. A score on a test the vendor wrote, scored and reported is the weakest form of evidence available, even when the vendor is acting in complete good faith.

This pattern is not unique to one company. The recent Gemini 3.7 Flash release in 2026 published its gains the same way, and the cross-vendor benchmark comparisons we looked at for GPT-5.6 show how quickly the picture changes once numbers are measured on a common harness.

What Did the Independent Evaluation Actually Find in 2026?

Independent measurement exists, and it tells a more sober story than the launch numbers. Artificial Analysis published its own evaluation of V4 Pro 0813 in 2026, placing it at 53 on its Intelligence Index at maximum reasoning effort. That index is a composite built from around ten separate benchmarks spanning reasoning, coding, agentic tool use and knowledge, run on a common harness across models.

Fifty-three is a solid mid-table result rather than a frontier one. The same index places Claude Opus 5 at 63, Claude Fable 5 at 62, GPT-5.6 Sol and Grok 4.6 at 61 and Kimi K3 at 60, with GLM-5.2 level with DeepSeek at 53. For a model at a fraction of frontier pricing, that is a genuinely strong value position, and it is a much more useful sentence than any preview-to-GA delta.

The number that should change a decision

The independent index separates V4 Pro 0813 from DeepSeek's own cheaper V4 Flash 0731 by a single point. At the post-change peak rates, Pro costs 3.96 US dollars per million output tokens against 1.32 for Flash. Before you budget for a Pro price rise in 2026, test whether Flash clears your quality bar, because the gap the independent evaluation measures is far narrower than the gap in price.

What Did DeepSeek Actually Publish, and What Is Still Missing in 2026?

More than the commentary suggests, and this is worth correcting because several write-ups in 2026 describe the release as entirely silent. DeepSeek published a news entry dated 13 August 2026 and a matching changelog entry, both covering the GA rollout, the agent improvements, the low, high and max thinking effort levels, native Responses API support optimised for Codex, and the upcoming pricing change. There is also a Hugging Face model card for DeepSeek-V4-Pro-0813 with open weights under an MIT licence.

What is thinner is the evaluation trail. The benchmark table in the announcement is published as an image rather than as text, which means the scores cannot be read by tooling, quoted precisely without transcription, or checked against a machine-readable record. The announcement carries no documented evaluation methodology, no description of the harness and no baseline conditions.

None of that is misconduct. Publishing a chart as an image is a formatting choice plenty of vendors make in 2026, and DeepSeek released the weights, the model card and a technical report, which is more openness than several better-funded competitors offer. It is a procurement consideration rather than a scandal: when the numbers live only in an image and the methodology is undocumented, there is less to hold the vendor to if performance shifts after you have built on it.

This is the same category of issue we flagged when GLM-5.3 shipped without its weights in 2026. The question is what documentation you could point at in six months if the model you deployed stopped performing the way it did at launch.

How Should Teams Evaluate a Cheap Model in 2026?

By measuring cost per completed task rather than cost per token, and by treating any published rate as a temporary condition. Per-token price is the input to a cost model, not the cost model itself. A model that is half the price but needs three attempts, longer prompts or heavier review is more expensive in practice, and that difference never shows up on a pricing page.

StepWhat to do in 2026Why it matters
Price the task, not the tokenMeasure total tokens and retries needed to finish one real unit of work end to endReasoning models can consume far more output tokens than a headline rate implies
Benchmark on your own workloadRun fifty to a hundred of your actual prompts against two or three candidatesPublished scores are measured on tasks that are not yours
Check the rate card's expiryLook explicitly for introductory, promotional or time-limited wording before committingTwo major vendors changed rate cards within months of each other in 2026
Read the change historyReview how often the vendor has revised pricing in the past twelve monthsPast volatility is the best available predictor of future volatility
Test the cheaper tier firstEvaluate the vendor's own smaller model before assuming the flagship is requiredThe independent index gap between Pro and Flash is one point in 2026
Plan an exitKeep prompts and evaluation harnesses portable across at least two providersA rate change with two days of notice is only survivable if switching is cheap

The expiry point generalises well beyond one vendor. Gemini 3.7 Flash launched in 2026 on an explicitly introductory rate that doubles on 1 January 2027, and DeepSeek is restructuring its card in August 2026. When two independent vendors reprice within months of each other, the reasonable conclusion is that cheap introductory AI pricing in 2026 is a market phase rather than a baseline, and budgets should be built on that assumption.

What Are the Common Mistakes in 2026?

The recurring errors are all versions of treating a launch-day figure as a durable fact. Each of these is avoidable with an hour of work, and each of them has cost teams real money in 2026.

Key Takeaways for 2026

The short version for anyone with a decision to make this week. DeepSeek V4 Pro remains a strong value proposition even after the change, and the argument here is about arithmetic and evidence quality rather than about avoiding the model.

If your model choice in 2026 was made on price, the work this week is small and specific: pull your last thirty days of token usage, split it by cache hit, cache miss and output, apply the peak and off-peak rates above, and see which side of your threshold the answer falls on. For a fuller picture of how the current model field compares, our breakdowns of what GPT-5.6 actually changed in 2026 and the Gemini 3.7 Flash release cover the same ground for the other major vendors.

DeepSeek V4 Pro 0813 in 2026: FAQs

When exactly does DeepSeek V4 Pro pricing change in 2026?

DeepSeek's own API pricing page states that the new prices take effect at 16:00 UTC on 16 August 2026. That is 21:30 IST on 16 August 2026 for teams in India, and midnight on 17 August in Beijing, which is why much of the coverage in 2026 reports the change as a 17 August event. From that moment the single published rate card is replaced by peak and off-peak billing, with off-peak rates set at half the peak rates.

Is the DeepSeek V4 Pro price increase really 12x in 2026?

The 12x figure is real but applies to one line of the rate card, not the whole bill. Cache-hit input moves from 0.003625 to 0.044 US dollars per million tokens at peak, which is roughly 12.1 times. Cache-miss input rises about 3 times and output about 4.6 times at peak. Off-peak rates are half of peak throughout. Most real workloads land between 1.8 and 3.7 times their current 2026 cost depending on cache behaviour and time of day.

What will DeepSeek V4 Pro cost after the 2026 price change?

Per DeepSeek's published pricing page, V4 Pro peak rates are 0.044 US dollars per million cache-hit input tokens, 1.32 per million cache-miss input tokens and 3.96 per million output tokens. Off-peak rates are exactly half, at 0.022, 0.66 and 1.98. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC in 2026, and every other hour is off-peak. Confirm current rates on the pricing page before you budget.

Are the DeepSeek V4 Pro 0813 benchmark scores independently verified in 2026?

Partly, and the two pictures differ in usefulness. The headline gains DeepSeek published, such as DeepSWE moving from 12.8 to 62.7, are self-comparisons against its own V4 Pro Preview rather than against any competitor. Separately, Artificial Analysis has published an independent evaluation in 2026 placing V4 Pro 0813 at 53 on its Intelligence Index at maximum reasoning effort, which is mid-table and one point above DeepSeek's own cheaper V4 Flash 0731.

Did DeepSeek launch V4 Pro 0813 without an announcement in 2026?

No, and reports saying so are inaccurate. DeepSeek published a news entry and a changelog entry dated 13 August 2026 covering the GA rollout, the agent improvements, the low, high and max thinking effort levels, native Responses API support and the upcoming pricing change. There is also a Hugging Face model card for DeepSeek-V4-Pro-0813 under an MIT licence. What the news post does not carry is a written benchmark table, since the scores are published only inside an image, or a documented evaluation methodology.

Should marketing teams switch to DeepSeek V4 Pro in 2026?

Not on the strength of the launch numbers alone. The defensible 2026 approach is to price the workload on the rate card that applies after 16 August, benchmark the model on your own tasks rather than on published scores, and compare it against DeepSeek V4 Flash as well as against other vendors, since the independent index separates Pro and Flash by a single point while Pro costs several times more per million output tokens at peak.

Re-running your AI model costs before the rate card moves?

Distk helps growth, marketing and founding teams across India and international markets build model selection on their own workloads and their own numbers, rather than on launch-day benchmarks. If you need a second pair of eyes on what a pricing change does to your stack, we can work through it with you.

Talk to Distk →