Skip to content
deepseekprice

Unofficial community resource. Not affiliated with DeepSeek.

Updated for the 16 August 2026, 16:00 UTC rate change

DeepSeek API pricing, in your timezone

DeepSeek is the only major API that charges by the hour of day. This page tracks the 50% off-peak window against your local clock, compares the new rate card with the old one meter by meter, and prices your real token volume against the alternatives.

DeepSeek runs at half price for the entire US workday. The peak windows land overnight across the Americas, so every mainland US timezone gets a continuous 15-hour off-peak run that covers 9-to-5 in full — in winter and summer alike. If you work US hours, the discounted rate is your default, not something you have to schedule for.

Live billing status

Checking your clock… $3.96 / M output tokens right now

DeepSeek bills on a UTC clock, so the discounted window lands at a different hour depending on where you are. Enable JavaScript to see the current state in your timezone.

--:--:--

until the rate changes

Peak output
$3.96/M
Off-peak output
$1.98/M

Today where you are

UTC UTC+00:00
00:0006:0012:0018:0024:00
  • Peak 01:00 – 04:00 UTC
  • Peak 06:00 – 10:00 UTC
  • Off-peak all remaining hours

17h of every day is discounted by 50%. Billing follows the UTC clock, not your local one — shifting a batch job by a couple of hours is often the cheapest optimisation available.

Summary of the price change

Before — flat, all hours

$0.87 /M

After — off-peak

$1.98 /M 2.3× the old rate

After — peak

$3.96 /M 4.6× the old rate

DeepSeek V4-Pro output, per million tokens. The old card was flat — one rate around the clock — so even the new off-peak rate is an increase on it.

New rates vs old rates

All figures are US dollars per million tokens. DeepSeek meters input twice — once at the cache-hit rate and once at the cache-miss rate — so the effective input price depends on how stable your prompt prefix is.

DeepSeek V4-Pro

Flagship model. Higher capability, lower concurrency allowance.

deepseek-v4-pro 1M context · 384K max output · 500 concurrency
DeepSeek V4-Pro price per million tokens, before and after the change. The old card was flat; the new one splits peak and off-peak hours.
Per 1M tokens Before the change From 16 Aug 2026 Increase
Flat, all hours Peak Off-peak at peak at off-peak
Input — cache hit Prompt prefix served from DeepSeek’s cache $0.003625 $0.044 $0.022 12× 6.1×
Input — cache miss Prompt processed fresh, no cache reuse $0.435 $1.32 $0.66 3.0× 1.5×
Output Generated tokens, including reasoning tokens $0.87 $3.96 $1.98 4.6× 2.3×

The old card had no time-of-day pricing — one rate applied around the clock. Both increase columns therefore compare against that single flat rate, and even the new off-peak rate sits above it.

DeepSeek V4-Flash

Cheaper, faster tier with a far higher concurrency allowance.

deepseek-v4-flash 1M context · 384K max output · 2,500 concurrency
DeepSeek V4-Flash price per million tokens, before and after the change. The old card was flat; the new one splits peak and off-peak hours.
Per 1M tokens Before the change From 16 Aug 2026 Increase
Flat, all hours Peak Off-peak at peak at off-peak
Input — cache hit Prompt prefix served from DeepSeek’s cache $0.0028 $0.014 $0.007 5.0× 2.5×
Input — cache miss Prompt processed fresh, no cache reuse $0.14 $0.44 $0.22 3.1× 1.6×
Output Generated tokens, including reasoning tokens $0.28 $1.32 $0.66 4.7× 2.4×

The old card had no time-of-day pricing — one rate applied around the clock. Both increase columns therefore compare against that single flat rate, and even the new off-peak rate sits above it.

DeepSeek API cost calculator

Enter what you actually send per day. The result blends the cache-hit and cache-miss rates, splits the day across peak and off-peak windows, and prices the same volume against the other major providers.

Start from a profile

80%

Share of input tokens DeepSeek can serve from cache. Coding assistants and agent loops re-send the same prefix constantly and often sit above 90% — DeepSeek Harness (dsh) sessions are a real-world example.

50%

Drag to zero to see what a fully off-peak schedule would cost.

From 16 Aug 2026

per month · per day

Before the change

per month · per day

Monthly difference

Running the same volume entirely in off-peak hours: per month, saving .

Same volume, other providers · monthly

  • DeepSeek V4-Pro
  • Anthropic Claude Sonnet 5
  • Anthropic Claude Opus 5
  • OpenAI GPT-5.6-sol
  • OpenAI GPT-5.6-terra
  • OpenAI GPT-5.6-luna
  • Google Gemini 3.1 Pro Preview

Comparisons assume the same token volumes and the same cache hit rate everywhere, which is a simplification: cache behaviour differs by provider. Only DeepSeek charges by time of day, so the other rows are flat.

Per-provider caveats
  • Anthropic Claude Sonnet 5 — Writing to the 5-minute cache costs $2.50 per million tokens on top of the read rate.
  • Anthropic Claude Opus 5 — Anthropic charges a separate premium for writing to the cache.
  • OpenAI GPT-5.6-sol — Caching is automatic; there is no separate cache-write charge.
  • OpenAI GPT-5.6-terra — Caching is automatic; there is no separate cache-write charge.
  • OpenAI GPT-5.6-luna — Caching is automatic; there is no separate cache-write charge.
  • Google Gemini 3.1 Pro Preview — Rates shown apply to prompts of 200K tokens or fewer; longer prompts move to a higher tier. Google charges hourly storage for explicit context caches on top of the read rate.

Price change timeline

New rates in

  1. 26 February 2025 background

    Off-peak discounts introduced

    DeepSeek first applies time-of-day pricing, discounting V3 by 50% and R1 by 75% during its quieter hours.

  2. 5 September 2025 background

    Off-peak discounts withdrawn

    The time-of-day mechanism is retired in favour of a permanent price cut, leaving a single flat rate around the clock.

  3. 13 August 2026

    New rates announced

    DeepSeek publishes a substantially higher rate card and brings time-of-day pricing back — this time as a peak surcharge rather than an off-peak discount on the old level.

    Official source ↗
  4. 16 August 2026, 16:00 UTC

    New rates take effect

    Requests are billed against the new card from this instant, and the peak/off-peak split starts applying on the UTC windows.

    Official source ↗

Effective date as published: 16 August 2026, 16:00 UTC.

Frequently asked questions

Why did DeepSeek raise its API prices?

DeepSeek published the new rate card on 2026-08-13. Peak output on DeepSeek V4-Pro goes from $0.87 to $3.96 per million tokens — 4.6× the old rate.

The company has not published a cost breakdown, so any explanation beyond the announcement itself is speculation. What the numbers show is a move away from the aggressive undercutting that defined DeepSeek’s earlier pricing, toward rates closer to the rest of the market.

When exactly do the new DeepSeek prices take effect?

The new rates apply from 16 August 2026, 16:00 UTC. Requests billed before that instant use the old card; requests after it use the new one.

Because billing runs on UTC rather than your local clock, the switch lands mid-working-day in some regions. The countdown at the top of this site tracks it in your own timezone.

What are DeepSeek’s off-peak hours in my timezone?

DeepSeek defines its windows on the UTC clock. Peak hours run 01:00–04:00 UTC and 06:00–10:00 UTC; every other minute of the day is off-peak, which works out to 17h of discounted time per day.

Where you are decides how convenient that is. In the mainland United States the windows fall overnight and in the evening, leaving a continuous off-peak run of at least 15 hours that covers the whole 9-to-5 working day in every US timezone, year round.

The table on this site converts the windows to whatever timezone your browser reports.

How much cheaper is the DeepSeek off-peak rate?

Off-peak rates are 50% lower than peak, applied to input and output alike. On DeepSeek V4-Pro that is $1.98 per million output tokens instead of $3.96.

It is worth being clear about what the discount is measured against: it is half of the new peak rate, not a return to the old one. Off-peak output still costs 2.3× what the same tokens cost under the previous flat card.

The rate is decided by when the request reaches DeepSeek, not when your job was queued locally. Batch work that can tolerate delay is the clearest way to capture it.

How does cache-hit pricing work?

Input tokens are billed at two different rates. Tokens DeepSeek can serve from its prompt cache cost $0.044 per million at peak; tokens it has to process fresh cost $1.32 per million — a gap of 30×.

Caching keys on identical prefixes, so it rewards keeping the stable part of a prompt — system instructions, tool definitions, a file you are iterating on — byte-identical at the front, with the varying part at the end. Reordering that prefix between calls throws the cache away.

Coding assistants and long agent loops re-send the same context repeatedly and commonly sit above 90% cache hits. That is why the calculator on this site defaults to 80% rather than assuming every token is billed at full price.

What cache hit rate should I actually expect?

It depends entirely on prompt shape, and the spread is enormous. A measured run of 613 requests from a coding agent — 155.9M of prompt tokens — came in at 98.1%, because an agent re-sends a long, stable context on every turn.

That is close to the ceiling, not the average. The first request of a session in the same run hit only 16.8% — there is nothing to hit yet — and workloads that put anything variable at the front of the prompt, such as a timestamp or a request ID, measure near zero however long they run.

Do not budget on someone else’s number, including this one. Both figures needed to compute your own are returned on every API response.

Does the price increase make caching less valuable?

Relatively yes, absolutely no — and the relative move is the surprising one. The cache-hit meter rose 6.1× against the old card, more than the cache-miss meter at 1.5× or output at 2.3×. The cheapest meter went up the most.

The consequence is counterintuitive: the better you cache, the harder this change lands. Repricing the measured 98.1% workload on the new off-peak card multiplies its bill by 2.7×. The identical tokens with no caching at all would have gone up only 1.5×.

In absolute terms caching is still worth more than it was, because it is a percentage off a bigger number. On that same run it removed $97.56 from a bill that would otherwise have been $104.39.

How do I measure my own cache hit rate?

Every API response carries the two counters in its `usage` object: `prompt_cache_hit_tokens` and `prompt_cache_miss_tokens`. They are disjoint and sum to the prompt size, so your hit rate is the first divided by the total. Log them and the guesswork ends.

If you use DeepSeek Harness, they are already on disk — it records a token breakdown for every request in `~/.dsh/sessions`. The audit script this site publishes at /dsh-cache-audit.py reads those logs and prints your real hit rate and bill. It has no dependencies and sends nothing anywhere.

Does DeepSeek charge extra to write to the cache?

No. Caching is automatic and there is no separate write charge — you are billed at the hit rate or the miss rate, and nothing else.

This is a real difference from some competitors rather than a technicality. Anthropic bills a premium for writing to its prompt cache, which has to be earned back through reads before caching pays for itself. On DeepSeek the first request simply costs the miss rate, as it would have anyway.

The trade-off is that you get no guarantee in return. DeepSeek describes its caching as best-effort, and unused entries are cleared after a few hours to a few days.

What is the difference between V4-Pro and V4-Flash?

Both carry a 1M token context window and the same maximum output length. They differ on price and on how many requests you may run at once: DeepSeek V4-Pro allows 500 concurrent requests, DeepSeek V4-Flash allows 2,500.

DeepSeek V4-Flash output costs $1.32 per million at peak against $3.96 for DeepSeek V4-Pro. Both models moved to the new card on the same date and share the same peak windows.

Is V4-Flash actually cheaper than V4-Pro in practice?

The rate card says 3.0× on output. Your bill may not agree, because price per token is only half of what you pay — the other half is how many tokens the model spends reaching an answer, and the smaller model does not always spend fewer. One published hands-on comparison ran both across three real coding tasks and found the cheaper model produced better work while consuming enough extra tokens to make the two bills roughly equal.

A second factor is usually larger than the model choice: reasoning effort. In the measured run on this site, 73.6% of all generated tokens were thinking tokens — invisible in the response, billed in full at the output rate, which is the most expensive meter on the card.

So the honest answer is that neither the headline discount nor a benchmark table settles it. Run your own workload on both, compare the total token counts rather than the rates, and check what reasoning effort you left it on.

Is DeepSeek still cheaper than Claude, OpenAI and Gemini after the increase?

For most workloads, yes — but by a smaller margin than before, and the gap narrows further against the cheaper tiers of each provider once you account for cache hits.

The right comparison depends on your own mix of input, output and cache rate, which is what the calculator is for: it applies your volumes to every rate card at once instead of comparing headline numbers that assume a workload you may not have.

Does the price change affect the DeepSeek chat app or only the API?

The rate card covers API usage, billed per token against your API account balance.

Consumer access through the DeepSeek app and website is a separate product on separate terms, and is not what the prices on this site describe.

Do credits bought at the old price keep their value?

Balance is held in currency, not in tokens. Credit already on your account stays valid, but from the change-over each request draws down that balance at the new per-token rates, so a given balance buys proportionally fewer tokens than it did before.

Check the official pricing page for the terms that actually govern your account: https://api-docs.deepseek.com/quick_start/pricing/

Figures are compiled from DeepSeek’s public announcements and pricing page and may lag behind changes. Verify against the official source before committing spend.