DeepSeek-V4-Pro and DeepSeek-V4-Flash bill at four different rates apiece, not two. Off-peak output on V4-Pro runs $1.98 per 1M tokens; peak output on the same model runs $3.96, exactly double, for the 7 hours a day (01:00 to 04:00 and 06:00 to 10:00 UTC) the peak window covers. V4-Flash follows the same 2x split at $0.66 off-peak and $1.32 peak.
Reasoning effort, cache hits and a workload's own timezone move the real bill further than the headline rate does.
Stop pricing from the $/1M row alone and check which 7 hours you're calling in.
The current rate card
Model | Cache-miss input | Cache-hit input | Output | Context | Max output |
|---|---|---|---|---|---|
deepseek-v4-flash, off-peak | $0.22 | $0.007 | $0.66 | 1M tokens | 384K tokens |
deepseek-v4-flash, peak | $0.44 | $0.014 | $1.32 | 1M tokens | 384K tokens |
deepseek-v4-pro, off-peak | $0.66 | $0.022 | $1.98 | 1M tokens | 384K tokens |
deepseek-v4-pro, peak | $1.32 | $0.044 | $3.96 | 1M tokens | 384K tokens |
Peak: 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, 7 hours a day. Off-peak: the remaining 17 hours. † deepseek-v4-flash-vision-exp prices identically to deepseek-v4-flash; see the images section below for its per-image token cap. All prices per 1M tokens, USD, collected 2026-09-09 from DeepSeek's API documentation and cross-checked against two independent trackers.
Both models list at the same off-peak input, cache-hit input and output ratio to each other, roughly 3x apart on every row.
V4-Flash off-peak input runs $0.22 against V4-Pro's $0.66. Output off-peak runs $0.66 against $1.98. The peak multiplier is identical across both models: exactly 2x on every column, every time.
Full specifications for the higher-priced tier live on the DeepSeek-V4-Pro model page.
Basis and method
This article aggregates published figures rather than running its own benchmark or load test. Every price above traces to the rate card on DeepSeek's own API pricing documentation, effective 2026-08-16 16:00 UTC. No figures on this page were produced by testing.
The timezone-overlap and blended-rate figures below are LLM Waves Research's own arithmetic on top of the published peak-window hours, not a number DeepSeek or any competitor publishes. Full source list and methodology: llmwaves.com/methodology.
This guide sorts by when you call the API, not by which model you call, since the peak window changes the bill by up to 2x regardless of model choice.

Takeaway: V4-Pro output costs $1.98 off-peak and $3.96 peak, a flat 2x gap that holds on V4-Flash too, at $0.66 and $1.32.
How much does the DeepSeek API cost per million tokens right now?
DeepSeek-V4-Pro costs $0.66 per 1M input tokens and $1.98 per 1M output tokens off-peak, $1.32 and $3.96 peak. DeepSeek-V4-Flash costs $0.22 and $0.66 off-peak, $0.44 and $1.32 peak.
These are the rates in DeepSeek's August 13, 2026 change-log entry, which took effect 2026-08-16 16:00 UTC and replaced a flat-rate card with no time-of-day split.
A page quoting $0.14 input or $0.28 output for either model, with no peak or off-peak split, is pricing DeepSeek before this change and is out of date by more than two ratio points on every row.
For how these rates stack up against OpenAI, Claude and Gemini, see the site's LLM API pricing comparison hub.
When are DeepSeek's peak hours, and do they hit your team?
Peak runs 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, 7 hours a day, every day. Converted to local business hours, exposure is not close to even.

Takeaway: A 9-to-5 workday in China or Singapore overlaps DeepSeek's peak window for 7 of its 9 hours; the same workday in US Pacific or US Eastern overlaps it for zero.
A team on US Pacific or US Eastern time calls the API entirely off-peak during business hours, since both UTC windows fall between roughly 6pm and 6am local.
A UK team picks up 2 hours of peak inside a 9-to-5 day, an EU Central team 3 hours, an India team about 4.5 hours, and an Australia East team about 4 hours.
A China or Singapore team sits inside peak pricing for nearly the entire working day, broken only by a 2-hour off-peak gap between the two UTC windows that lands over the local lunch break.
What would an always-on workload actually pay?
A workload that calls the API at a constant rate around the clock does not pay the off-peak rate or the peak rate. It pays a blend of both, weighted by how many of the 24 hours fall in each window.

Takeaway: V4-Pro output blends to $2.56 per 1M tokens for a uniform 24/7 workload, against $1.98 for US-business-hours-only traffic and $3.52 for traffic concentrated in a China or Singapore workday, a 78% spread driven entirely by call timing.
The always-on figure: (7 peak hours × $3.96 plus 17 off-peak hours × $1.98) divided by 24 equals $2.56 per 1M output tokens. A workload confined to 13:00 to 21:00 UTC, US business hours, never touches either peak window and pays the flat $1.98 off-peak rate the whole day.
A workload confined to 01:00 to 10:00 UTC, a China or Singapore business day, spends 7 of its 9 hours in peak and blends to $3.52, within 11% of the pure peak rate.
None of these three blended figures appear on DeepSeek's own pricing page; each follows directly from multiplying the published peak-window hours against the published rates.
Does thinking effort (low, high, max) change what you pay?
Yes, but not through a separate price. DeepSeek's V4-Pro general-availability announcement documents a reasoning effort setting, low, high or max, described as low for simple tasks, high for daily agent workflows, and max for complex tasks.
All three effort levels bill at the same per-token output rate. What changes is the volume of output tokens a call produces: a higher effort setting generates a longer reasoning trace before the visible answer, and every token of that trace bills as a normal output token.
A max-effort call on a task that a low-effort call would answer in 200 tokens can run several times that in reasoning tokens alone, at the same $1.98 or $3.96 per 1M rate, so effort moves the bill through token count, not through a rate multiplier the way OpenAI's o-series prices reasoning effort directly.
Budgeting for max effort on every call without checking whether the task needs it is the single fastest way to overshoot a token estimate built on low-effort numbers.
For how V4-Pro's price and accuracy compare directly against Claude's tiers rather than against a per-token rate alone, see DeepSeek vs Claude.
Is there a DeepSeek API free tier, or only a signup grant?
Unresolved as a standing offer. A widely repeated claim puts a 5M-token signup grant on every new API account; DeepSeek's current API documentation does not publish a universal grant at that or any other size, and no expiry window for one is documented.
Any promotional balance that does appear is account-, region-, or campaign-specific rather than a standing policy, and the free DeepSeek chat product is not evidence that API calls are free, since the two run on separate billing.
Checking the account's own billing dashboard is the only way to confirm what a given account actually has, rather than trusting either the "5M free tokens" or the "no free tier" claim found across the ranking pages for this keyword, since both circulate and neither is confirmed as DeepSeek's current, published policy.
What about images, batch jobs and reselling markups?
Three smaller mechanics change specific bills without changing the rate card.
DeepSeek-V4-Flash-Vision-Exp, the vision-capable variant, prices identically to standard V4-Flash.
Images are billed as input tokens with a hard cap of 384 tokens per image regardless of resolution, meaning there is no higher-resolution, higher-cost mode to opt into; a receipt or screenshot OCR workload pays the same per-image ceiling whether the source image is 200KB or 20MB.
A cache hit requires an exact match against a previously cached prefix; changing any earlier part of a prompt breaks the cache for everything after that point, and DeepSeek's context-caching documentation states cached entries clear automatically "within a few hours to a few days" with no guaranteed retention window. Caching is on by default and needs no code change to use.

Takeaway: Off-peak cache-hit input runs $0.007 on V4-Flash and $0.022 on V4-Pro, both roughly 30 times cheaper than the equivalent cache-miss rate, so a broken prefix match costs far more than the peak/off-peak swing does.
Concurrency, not a requests-per-minute ceiling, is what caps a DeepSeek account. Per DeepSeek's rate-limit documentation, the default is 500 concurrent requests on V4-Pro and 2,500 on V4-Flash and V4-Flash-Vision-Exp, and an account can request a capacity expansion above that default at no additional cost, granted against stated business need rather than automatically.
Reselling markups run in both directions. OpenRouter lists DeepSeek-V4-Pro at a flat $0.87 input and $1.74 output, no peak or off-peak split of its own.

Takeaway: OpenRouter's $0.87/$1.74 flat rate undercuts DeepSeek's own peak price by roughly a third but costs 32% more than DeepSeek's own off-peak input, so which side wins depends entirely on how peak-heavy the calling workload is.
What this article did not measure
This article did not run a latency test, a throughput test, or an output-quality benchmark against any model discussed.
Every figure above is a price, a limit, or a billing mechanic collected from vendor documentation and cross-checked trackers, not a performance measurement LLM Waves Research produced itself.
Usage of Deepseek as per Work Nature
Developers
Situational. Off-peak V4-Flash at $0.22/$0.66 is cheap enough for iteration, but a US-hours-only workload never touches peak pricing at all, making the peak/off-peak split irrelevant to most solo development schedules.
Startups
Best value on V4-Flash off-peak, situational once volume grows past a single developer's schedule and starts running around the clock.
Enterprise
Situational. The always-on blended rate of $2.56 per 1M output tokens on V4-Pro is the number to budget against, not the $1.98 headline off-peak figure, for any workload that does not deliberately schedule around the peak window.
High volume
Top pick for scheduling flexibility: a queue-based workload that can defer non-urgent calls into the 17-hour off-peak window captures the full 2x saving that a synchronous, always-on workload cannot.
Low latency
Insufficient data. No latency or throughput figures were collected in this pass.
On-device
Not applicable. DeepSeek's API models are not distributed for on-device inference under this pricing structure.
Self-hosting
Not applicable to this article. DeepSeek's open-weight models are a separate cost question from its hosted API pricing and are not covered here.
Non-English
Insufficient data. No non-English accuracy or cost figures were collected in this pass.
Is DeepSeek free?
No, not as a standing API policy. The DeepSeek chat web product is free to use; the API is metered per token at the rates above, and no confirmed universal free-tier grant exists in DeepSeek's current documentation.
Is the DeepSeek API cheaper than OpenAI or Claude?
On the published per-token rate, yes, DeepSeek's off-peak pricing undercuts both by a wide margin; at peak, or blended for an always-on workload, the gap narrows sharply and is workload-dependent rather than a fixed multiple, since neither OpenAI nor Claude runs a time-of-day pricing split to compare against.
What is DeepSeek's context window and maximum output?
Both current models list a 1M-token context window and a 384K-token maximum output per request.
What are DeepSeek's rate limits?
Concurrency-based: 500 concurrent requests on V4-Pro, 2,500 on V4-Flash and V4-Flash-Vision-Exp, with a free capacity expansion available on request. Exceeding the limit returns an HTTP 429 error.
What happened to deepseek-chat and deepseek-reasoner?
Both legacy aliases were retired on 2026-07-24, per an April 2026 change-log notice giving three months' warning.
They had pointed to the non-thinking and thinking modes of V4-Flash respectively; any integration still calling either alias by name has been failing or falling back since that date.
The numbers on this page describe two specific days in September 2026
Peak and off-peak pricing has already changed once this year. Check DeepSeek's own pricing documentation, not this article, before committing a production budget to any figure above.
