AI Model Guide

DeepSeek V4.1 Flash Pricing in 2026: The Rate Card, the Time Zone Problem and the Real Cost

The cheapest capable rate card of September 2026 has three prices per token type, two clocks, and a footnote that reverses an announcement. This is the budget owner's read: every number, translated to IST, compared to the frontier, and worked through on an illustrative workload.

Distk Editorial Sep 2026 13 min read

DeepSeek V4.1 Flash is billed per million tokens at three rates that each halve off-peak: cache-hit input 0.006 US dollars peak and 0.003 off-peak, cache-miss input 0.30 peak and 0.15 off-peak, output 1.20 peak and 0.60 off-peak. Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, which is 06:30 to 09:30 and 11:30 to 15:30 IST, so most of an Indian working day is peak. Against DeepSeek V4 Pro's own rate card, V4.1 Flash is about 3.3 times cheaper on output and about 7 times cheaper on cache hits. Against Claude Fable 5.1 and GPT-6 Astra at 10 and 50 US dollars, it is roughly 33 times cheaper on cache-miss input and 42 times cheaper on output at peak. The reason cache hits are near free is the model's KV cache, about a quarter the size of V4 Flash. New rates took effect at 04:00 UTC on 10 September 2026.

What Does DeepSeek V4.1 Flash Cost in 2026?

DeepSeek V4.1 Flash costs 0.30 US dollars per million cache-miss input tokens, 0.006 per million cache-hit input tokens and 1.20 per million output tokens during peak hours in 2026, with every rate halved off-peak to 0.15, 0.003 and 0.60. The rates are published on DeepSeek's API pricing page, took effect at 04:00 UTC on 10 September 2026, and apply to the model name deepseek-flash and to the retired legacy names that now route to it. Billing is deducted from a prepaid balance, and DeepSeek states that prices may change.

Token type (per 1M tokens)DeepSeek V4.1 Flash, peakDeepSeek V4.1 Flash, off-peakDeepSeek V4 Pro 0813, peakDeepSeek V4 Pro 0813, off-peak
Input, cache hit0.006 USD0.003 USD0.044 USD0.022 USD
Input, cache miss0.30 USD0.15 USD1.32 USD0.66 USD
Output1.20 USD0.60 USD3.96 USD1.98 USD
Concurrency limit2,500500

Three ratios are worth memorising. Cache-miss input costs 50 times cache-hit input. Output costs 4 times cache-miss input. And off-peak is exactly half of peak on every line. The V4.1 Flash overview explains the architecture; this guide is only about what the meter reads.

Why Is Cache-Hit Input Almost Free on DeepSeek V4.1 Flash in 2026?

Because the thing being cached got four times smaller. A cache hit means the model is re-reading input it has already processed, such as a system prompt, a set of tool definitions, or the earlier turns of an agent's session. What the provider stores to make that possible is the KV cache, and DeepSeek V4.1 Flash's Causal Encoder-Decoder design brings that down to about 890 bytes per token, roughly a quarter of DeepSeek V4 Flash. DeepSeek's announcement says the cache needs a quarter of the GPU memory and an eighth of the SSD of the previous generation.

DeepSeek states plainly that cache-hit charges often account for a large share of agent costs, and that compressing the cache cuts those costs. This is the same economic lever Anthropic pulled the same month when it cut Claude Fable 5.1 cache reads by 75 percent to 0.25 US dollars per million. DeepSeek's cache-hit rate is 0.006 at peak. The two vendors are pricing the same phenomenon, one at 40 times the other.

When Are DeepSeek's Peak Hours in Indian Standard Time in 2026?

DeepSeek's peak hours are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, Monday through Friday, with all other hours off-peak. In Indian Standard Time that is 06:30 to 09:30 and 11:30 to 15:30 on weekdays. The gap between 09:30 and 11:30 IST is off-peak, everything after 15:30 IST is off-peak, and weekends are off-peak throughout.

Window (weekdays)UTCISTRateWhat it means for an Indian team
Early morning01:00 to 04:0006:30 to 09:30PeakOvernight batch jobs must finish before 06:30 IST to stay off-peak.
Mid-morning gap04:00 to 06:0009:30 to 11:30Off-peakA two-hour half-price window inside the working day.
Midday06:00 to 10:0011:30 to 15:30PeakThe core of the Indian working day is full price.
Afternoon onward10:00 to 01:0015:30 to 06:30Off-peakLate afternoon, evening and night are half price.
Saturday and SundayAll dayAll dayOff-peakWeekend batch work is always half price.

For a US East Coast team, peak is 21:00 to 00:00 and 02:00 to 06:00 Eastern, meaning almost the entire US working day is off-peak. For a UK team, peak is 02:00 to 05:00 and 07:00 to 11:00 in summer time, so the afternoon is off-peak. The rate card is time-zone neutral in its wording and time-zone unequal in its effect, and Indian teams sit on the expensive side of that.

Scheduling rule for 2026

Split workloads into interactive and flexible. Interactive work, such as a chatbot answering a customer or an agent a marketer is watching, runs whenever it runs. Flexible work, such as nightly classification, bulk translation, content generation and report assembly, should be queued to start after 15:30 IST or before 06:30 IST, or on weekends. DeepSeek says this in one line: schedule flexible workloads off-peak to save. For an Indian team the saving is exactly half.

How Does DeepSeek V4.1 Flash Compare to V4 Pro on Cost in 2026?

V4.1 Flash is cheaper than V4 Pro on every line of the rate card: about 7.3 times cheaper on cache hits, 4.4 times cheaper on cache-miss input, and 3.3 times cheaper on output, at both peak and off-peak. It also has five times the concurrency limit. DeepSeek states that third-party tests put V4.1 Flash ahead of V4 Pro on performance, cost, speed and total runtime, and its own benchmarks place V4.1 Flash above V4 Pro on most agentic rows.

The complication is the retirement that was announced and then walked back. On 10 September DeepSeek said V4 Pro requests would route to V4.1 Flash at V4.1 Flash rates from 14 September. On 11 September the pricing page added a footnote continuing V4 Pro service with billing unchanged. Either way, a team paying V4 Pro rates in 2026 for work that V4.1 Flash does as well is paying three to seven times more than it needs to. The migration guide covers how to test that safely; the August V4 Pro rate change guide covers how the Pro rate card got where it is.

How Does DeepSeek V4.1 Flash Compare to GPT-6 Astra, Claude Fable 5.1 and Gemini in 2026?

At peak, DeepSeek V4.1 Flash is roughly 33 times cheaper than Claude Fable 5.1 and GPT-6 Astra on cache-miss input and about 42 times cheaper on output. Against Gemini 3.5 Flash-Lite, Google's cheapest published tier, it is level on cache-miss input at peak and about half the price off-peak, and about half the output price at peak. The table uses each vendor's published list rates as of mid-September 2026; the full cross-vendor picture, including fit by workflow, is in the September 2026 model pricing comparison.

Model (published list rate, per 1M tokens)Cache-miss inputCache-hit or cache-read inputOutput
DeepSeek V4.1 Flash, peak0.30 USD0.006 USD1.20 USD
DeepSeek V4.1 Flash, off-peak0.15 USD0.003 USD0.60 USD
DeepSeek V4 Pro, peak1.32 USD0.044 USD3.96 USD
Gemini 3.5 Flash-Lite0.30 USDNot stated on model page2.50 USD
Gemini 3.8 FlashNot published on model pageNot publishedNot published
Claude Fable 5.110.00 USD0.25 USD (cache read)50.00 USD
GPT-6 Astra (Standard)10.00 USDSeparate rate, not enumerated at launch50.00 USD

A price gap this large changes the question. The question is no longer whether DeepSeek V4.1 Flash is as good as a frontier model, because on most of DeepSeek's own agentic benchmarks it trails the newest flagships on the hardest tasks, as the benchmark guide shows. The question is which of your workflows are good enough at a fortieth of the price, and whether the data-handling and availability profile fits. Those are workflow questions, and they belong in an evaluation set, not a rate card.

A Worked Cost Illustration for 2026

The table applies only the published rates to three hypothetical monthly workloads. The token volumes are illustrative placeholders chosen to show the shape of the change, not measured figures from any deployment. Each workload is shown at DeepSeek V4.1 Flash peak, V4.1 Flash off-peak, V4 Pro peak, and Claude Fable 5.1 list rates.

Illustrative monthly workloadV4.1 Flash, all peakV4.1 Flash, all off-peakV4 Pro, all peakClaude Fable 5.1
Classification: 40M cache-miss in, 20M cache-hit in, 4M out16.92 USD8.46 USD69.52 USD605.00 USD
Content pipeline: 30M cache-miss in, 60M cache-hit in, 30M out45.36 USD22.68 USD161.04 USD1,815.00 USD
Agent sessions: 20M cache-miss in, 400M cache-hit in, 40M out56.40 USD28.20 USD202.40 USD2,300.00 USD
Clearly labelled as an illustration

The costs above are arithmetic applied to published per-token rates at invented volumes. Claude Fable 5.1 is calculated at 10 USD cache-miss input, 0.25 USD cache read and 50 USD output. They exclude thinking-token overhead, retries, platform fees and any rate changes after 13 September 2026. They are not a forecast or an observed cost. Substitute your own measured token counts before using any of it in a budget.

The agent-session row is the one DeepSeek's architecture is built for. Four hundred million cache-hit tokens is what a fleet of long-running agents re-reading their context produces, and at 0.006 or 0.003 US dollars per million that line item is 2.40 or 1.20 US dollars. On a frontier model with 0.25 cache reads the same line is 100 US dollars, and on a model that does not discount cache reads at all it would be the largest line on the invoice.

What Does Reasoning Effort Do to the Bill in 2026?

DeepSeek V4.1 Flash runs in thinking mode by default and exposes reasoning effort as an integer from 1 to 100. Thinking produces tokens, and output tokens are the most expensive line on the rate card at 1.20 US dollars per million at peak. DeepSeek's published benchmark scores use effort 100. A classification or extraction workflow does not need effort 100, and running it there is paying for reasoning the task does not use. The discipline in 2026 is the same as for any model with an effort dial: set it per workflow, measure accuracy at each level against your own evaluation set, and let the lowest level that passes be the default.

What Are the Common Mistakes With DeepSeek Pricing in 2026?

Key Takeaways for 2026

Distk helps growth teams across India and internationally instrument model spend per workflow, schedule flexible work into off-peak windows, and decide which tasks earn a frontier rate and which do not. If DeepSeek's September 2026 rate card is on your desk, that instrumentation is where we start.

DeepSeek V4.1 Flash Pricing in 2026: FAQs

How much does DeepSeek V4.1 Flash cost per million tokens?

At peak: cache-hit input 0.006 USD, cache-miss input 0.30 USD, output 1.20 USD. Off-peak every rate halves: 0.003, 0.15 and 0.60. Rates effective 04:00 UTC on 10 September 2026.

What are DeepSeek's peak hours in IST?

06:30 to 09:30 and 11:30 to 15:30 IST, Monday to Friday. That is 01:00 to 04:00 and 06:00 to 10:00 UTC. All other hours and all weekend hours are off-peak at half price.

Is DeepSeek V4.1 Flash cheaper than V4 Pro?

Yes on every line: about 7 times cheaper on cache hits, 4.4 times on cache-miss input and 3.3 times on output, with five times the concurrency limit. DeepSeek says third-party tests also found it faster.

How does DeepSeek V4.1 Flash compare to GPT-6 Astra and Claude Fable 5.1 on price?

Both flagships list at 10 USD input and 50 USD output per million. At peak, V4.1 Flash is roughly 33 times cheaper on cache-miss input and 42 times cheaper on output. Claude's 0.25 USD cache read is about 40 times DeepSeek's 0.006 cache hit.

Why are DeepSeek cache hits so cheap?

The KV cache that makes a cache hit possible is about 890 bytes per token on V4.1 Flash, roughly a quarter of V4 Flash, so DeepSeek can hold far more cached context per unit of hardware and price it accordingly.

Does reasoning effort change the price?

Indirectly. Thinking generates output tokens, billed at 1.20 USD per million at peak. Effort runs from 1 to 100; benchmarks use 100. Set it per workflow to avoid paying for unneeded reasoning.

Model the bill at peak, then earn the off-peak discount

Distk instruments token spend per workflow, separates interactive from flexible work, and schedules the flexible half into DeepSeek's off-peak windows. If your 2026 AI budget has a DeepSeek line, we make sure it is the right size.

Start the conversation →