DeepSeek's Price Hike Gets Real: Up to 4.7x on Output
DeepSeek's Price Hike Gets Real: Up to 4.7x on Output
Last week I ended the DeepSeek price-hike post with a promise: I'd follow up the moment the numbers dropped. They're out — and on August 16 at 16:00 UTC, the old rate card is gone. V4 Flash output goes from $0.28 to $0.66 per million tokens off-peak — and to $1.32 at peak, 4.7x what it costs today. And here's the twist the July announcement didn't prepare anyone for: off-peak went up too. Everything goes up; peak hours just double it again.
V4 Flash launched April 24 at $0.14 input / $0.28 output and hasn't moved since. V4 Pro's 75% launch discount became permanent on May 22, landing at $0.435 / $0.87. On August 6 DeepSeek posted a notice calling the coming increase "significant" without naming a number. The number is now on the official pricing page, and it's bigger than the version DeepSeek floated in July.
My own coding agent runs on V4 Flash, so I have skin in this one — and I've been re-reading my usage with the new card in hand.
| Line item (per 1M tokens) | Today | Off-peak (Aug 16) | Peak (Aug 16) |
|---|---|---|---|
| V4 Flash — cache hit | $0.0028 | $0.007 | $0.014 |
| V4 Flash — cache miss | $0.14 | $0.22 | $0.44 |
| V4 Flash — output | $0.28 | $0.66 | $1.32 |
| V4 Pro — cache hit | $0.003625 | $0.022 | $0.044 |
| V4 Pro — cache miss | $0.435 | $0.66 | $1.32 |
| V4 Pro — output | $0.87 | $1.98 | $3.96 |
Official numbers from the DeepSeek pricing page, retrieved August 13. Peak hours: 01:00–04:00 and 06:00–10:00 UTC.
Every line item moves on August 16. Off-peak is 1.5–6.1x today's price, and peak doubles that again — log scale, official numbers. Chart by the author.
The off-peak surprise
Back in July, when the peak-hour plan surfaced and then got confirmed, the framing was simple: prices double in the busy windows, and off-peak, nothing moves. TheNextWeb's write-up even spelled it out — Pro output would go from ¥6 to ¥12 per million at peak, about $0.85 to $1.70. What actually lands is different.
Off-peak Pro output is $1.98, not $0.87. Peak Pro output is $3.96, not $1.70. On the output side, the final card is two to two-and-a-half times the July plan.
Put another way: if you run only in off-peak hours — the "safe" hours everyone was told would stay unchanged — Pro output now costs 2.3x what you pay today, and Flash output 2.4x. The peak surcharge is the headline, but the off-peak hike is the body of the bill.
The cache-hit discount took the biggest beating
Prefix caching was quietly the best deal in AI. Feed a long system prompt, and every token after the first request costs a fraction. Flash cached input today is $0.0028 — 2% of the miss price. Pro cached input is $0.003625 — 0.83%, essentially free. That's how long-context agents and RAG apps kept their bills small.
The new card changes the arithmetic. Pro cache hits go to $0.022 off-peak and $0.044 at peak — 6.1x and 12.1x today's rate, and now 3.3% of the miss price instead of 0.83%. Flash cache hits go to $0.007 / $0.014 — 2.5x and 5x.
Caching still saves money, dramatically. But the "essentially free" era for long prompts is over.
What peak-hour pricing does to each line item versus today. The two cache-hit lines are the shockers — 5x and 12.1x. Chart by the author.
What it does to an actual bill
Abstract multipliers are easy to shrug off, so let me put a bill on it. Take a modest coding-agent workload: 20 million cache-miss input tokens and 4 million output tokens a week — roughly a small team running an agent loop on V4.
| Weekly bill (20M in, 4M out) | V4 Flash | V4 Pro |
|---|---|---|
| Today | $3.92 | $12.18 |
| Off-peak only | $7.04 | $21.12 |
| Peak only | $14.08 | $42.24 |
Illustrative bill, cache-miss input, official rates. Your mix will differ — cache hits soften it, thinking mode worsens it.
If you run around the clock, the blended rate lands between the two: peak is 7 hours of the day, off-peak 17. Spread evenly, Flash output averages about $0.85 per million (3x today) and Pro output about $2.56 (2.9x).
There's a structural shift here too: output now costs exactly 3x input on both models. It was 2x. Thinking mode burns output tokens by the truckload, and DeepSeek is now pricing that directly.
When is peak, exactly?
The docs give two windows: 01:00–04:00 and 06:00–10:00 UTC. That's 09:00–12:00 and 14:00–18:00 in Beijing, and 10:00–13:00 and 15:00–19:00 in Seoul. Seven hours a day, every day — the page lists no weekend exemption, so don't count on one.
The two 2x windows, in UTC, Beijing and Seoul time. Off-peak is the other 17 hours — cheaper than peak, but no longer today's price. Chart by the author.
For the record on the start time: the new card takes effect 16:00 UTC on August 16 — Sunday afternoon in the US, Sunday evening in Europe, and 1 a.m. Monday in Seoul. Not the friendliest hour to flip the switch, but the docs are unambiguous.
Is it still the cheapest? Mostly.
The honest answer is yes, with an asterisk. Stack the new output prices against the list prices from my earlier comparisons and the gap is still absurd — just smaller.
| Output per 1M tokens | Price |
|---|---|
| V4 Flash — today | $0.28 |
| V4 Flash — off-peak | $0.66 |
| V4 Pro — today | $0.87 |
| V4 Flash — peak | $1.32 |
| V4 Pro — off-peak | $1.98 |
| V4 Pro — peak | $3.96 |
| Kimi K3 | $15 |
| Claude Opus 5 | $25 |
| GPT-5.6 Sol | $30 |
| Fable 5 | $50 |
List prices, August 2026. Frontier prices from official pages and press coverage; DeepSeek numbers from the new rate card.
Even at peak rates, DeepSeek sits 4–13x below the big four on output. Flash's launch price — the $0.28 bar — sat 50–180x below the same list prices. Chart by the author.
The asterisk: OpenAI's own budget tier, GPT-5.6 Luna, lists at $0.20 input / $1.20 output after last month's permanent cut. During peak hours, Flash output at $1.32 actually costs more than Luna's. The title of "cheapest serious model on earth" now comes with a clock attached — off-peak, DeepSeek still holds it; at peak, it's a race.
And the caveat I repeat every time pricing comes up: token price is a lying unit. Thinking models can burn several times more output per answer, which is exactly why the output-to-input ratio just went from 2:1 to 3:1. When I checked the Artificial Analysis numbers last week, Flash was still roughly 100x cheaper per test run than Claude. An increase like this trims that lead; it doesn't come close to closing it.
So what do you actually do?
The first move is scheduling: batch whatever you can into the 17 off-peak hours, because the difference is now real money, not pocket change. The second is caching — the discount shrank, but a 96%+ discount on repeated prefixes is still a 96%+ discount. If you haven't already, make sure your system prompt is prefix-friendly. The third is output tokens: thinking mode is the expensive part now, so tune reasoning effort like a budget line, because it is one.
And if you were deciding whether to build on V4: the price is still 4–13x below the big four on output — just not 50–180x anymore. Decide like a business, not like a lottery winner.
Why now (and why it isn't just greed)
I made this case last week, and the new numbers don't change it. One model burning roughly eight trillion tokens a day — more than every model on OpenRouter combined, per OpenCode's own disclosure — is a cost problem, not a growth story. DeepSeek's founder pre-announced roughly a 50% hike at a July investor briefing, and the company is raising a Series B of around RMB 50 billion (~$7B) at a pre-money valuation near RMB 500 billion, with signing targeted for late August. You do not walk into that meeting selling tokens 100x below market.
Time-of-day pricing is also just honest grid management. GPUs idle at 3 a.m. and choke at 3 p.m.; charging by the clock is the oldest trick in the economics book for smoothing exactly that kind of load. The off-peak hike on top of it is the part that isn't grid management — it's the subsidy ending.
My take
The verdict from last week hasn't changed: DeepSeek is still the best deal in frontier AI by a mile, and it just got more expensive on every line of the invoice. The gap to the big four on output is now "4–13x" instead of the "50–180x" Flash's launch price sat at. The price war isn't over — it's being wound down, from the inside. If anything changes before the 16th, I'll update this. And I'll be watching my own bill on the 17th.
Sources: DeepSeek Models & Pricing page (official, retrieved August 13, 2026) · DeepSeek API changelog · TheNextWeb, July 1, 2026 (peak-hour plan and July pricing) · Bloomberg, August 6, 2026 (increase notice; Kimi K3 and Fable 5 list prices) · deepseek.ai pricing reference (May 22 permanent Pro discount) · earlier posts on this blog: the August 7 price-hike analysis and the V4 Pro launch post. All prices are list prices in USD per 1M tokens unless noted; performance and pricing claims are vendor-reported. This blog is independent — no company mentioned here paid for coverage.
Comments
Post a Comment