Skip to main content

The Cheapest Model on Earth Just Announced a Price Hike. 8 Trillion Tokens a Day Will Do That.

The Cheapest Model on Earth Just Announced a Price Hike

AI · Pricing · DeepSeek

The Cheapest Model on Earth Just Announced a Price Hike. 8 Trillion Tokens a Day Will Do That.

Real talk: for the last few months, my coding agent has been running on a model that costs $0.14 per million input tokens. Not a typo. The same class of model from Anthropic costs $25 per million output tokens. DeepSeek V4 Flash showed up in April, looked at the price war, and basically won it by being 50-100x cheaper than everyone else.

And then, on August 6, DeepSeek dropped this notice:

"We plan to implement a comprehensive price increase for DeepSeek API services in the near term, with a significant expected rise. Please adjust your usage accordingly. The specific plan will be subject to the official notice."

"Significant." Not "modest." Not "slight readjustment." Significant.

Honestly? This was the most predictable plot twist of 2026. Here's why.

The number that broke the camel

DeepSeek didn't hike prices because they're greedy. They hiked prices because one model is eating the entire internet's token budget.

Look at these numbers, from OpenCode's own disclosure:

  • DeepSeek V4 Flash via OpenCode alone: 8 trillion tokens per day — 5T from free quotas, 3T from paid plans
  • All 400+ models combined on OpenRouter: ~6.6 trillion tokens per day

Let me say that again. One model, through one access point, consumes more tokens in a day than every model on OpenRouter combined. OpenRouter hosts Gemini, Claude, GPT, Llama, Mistral — 400+ of them. Doesn't matter. V4 Flash is the #1 token consumer on OpenRouter's weekly chart at 7.22T tokens, and on Vercel it's doing ~5.3T a week.

One model vs the whole routing platform

One model out-consumes an entire routing platform. Sources: OpenCode disclosure, OpenRouter, Vercel.

That's not a usage spike. That's a firehose pointed at one door.

How we got here (it was never going to stay this cheap)
DateEvent
Apr 24V4 Flash + V4 Pro launch. Flash at $0.14/$0.28, Pro at $0.435/$0.87, both with 1M context
May 22DeepSeek makes a 75% launch discount permanent — the "AI price war" headlines peak
mid-JulyPeak/off-peak pricing goes live: 2x price on weekday peaks (9-12, 14-18 Beijing time)
Jul 31V4-Flash-0731 official API beta — same architecture, retrained, notably better at coding/agent tasks
Aug 6"Comprehensive price increase... significant expected rise" — numbers TBD
Timeline of the drama

Three months of pricing drama, one chart.

See the pattern? Every time usage exploded, the response was more usage — cheaper. Until the math stopped working.

The boss basically told us this was coming

At an investor briefing in July, DeepSeek founder Liang Wenfeng laid out the pricing philosophy pretty plainly:

  • API pricing is built around a "reasonable profit model" — hardware costs should be recovered within 10 months of purchase
  • And the killer line: demand at current prices has "almost no elasticity""even if prices were raised by another 50%, token consumption would barely change"

Read that again. They pre-announced a 50% hike in July. The August 6 notice is just the formal version.

And the business context makes it make sense: ARR is already $400-500M, the V4 line has gross margin above 50%, and DeepSeek is raising a Series B of RMB 50 billion (~$7B) at a pre-money valuation of ~RMB 500B, with signing targeted for late August. When you're about to raise $7 billion, you don't go into the meeting with prices 100x below the competition.

The dry run: peak/off-peak pricing

The mid-July "time-of-day" pricing was the appetizer — a 2x surcharge on weekday peaks (9am-12pm, 2pm-6pm Beijing time). Goldman Sachs calculated the blended effect: about $0.35/M for V4 Pro and $0.12/M for Flash after weighting.

When your bill doubles

Weekday peaks, in Beijing and Seoul time.

For Korean users that means: 10am-1pm and 3pm-7pm KST = 2x. Shift batch work to night, and you're basically unaffected.

Even after the hike, it'll still be the cheapest thing on the market

Here's the part that should calm everyone down. Artificial Analysis weights cost per test run (because token price alone lies — a model that needs 10x more tokens to answer isn't actually cheap):

  • DeepSeek V4 Flash: $0.03 per test
  • Kimi K3: $0.86
  • GPT-5.6 Sol: $1.86
  • Claude Fable 5: $3.15
Still the cheapest

Weighted cost per test run, log scale. Source: Artificial Analysis (Aug 2026).

V4 Flash is 100x cheaper per test than Claude. A "significant" increase could mean doubling — and Flash would still be ~50x cheaper. Even a 3x hike keeps it the best deal in AI by a mile.

What might "significant" actually look like? Nobody knows yet (that's the honest part — the official numbers haven't dropped). Pure speculation on my part, but for context, here's what Flash pricing would look like at various bumps:

ScenarioInput $/MOutput $/MCached $/M
Today (off-peak)$0.14$0.28$0.0028
+30%~$0.18~$0.36~$0.004
+50%~$0.21~$0.42~$0.004
2x$0.28$0.56~$0.006

My guess: they won't go full 2x on the base rate — the 2x already exists as the peak mechanism. A base bump in the +30-60% range, possibly paired with changes to free quotas, feels most likely. But again: this is a guess. The notice says specifics are coming.

What I'm actually doing
  1. Moving batch work off peak hours. The 2x peak pricing already exists and it's cheap to just... schedule around it.
  2. Leaning hard into prompt caching. Cache-hit input is $0.0028 — 98% off. Long agent sessions that reuse context are already nearly free; that advantage survived the hike announcement, and I'd bet it survives the hike.
  3. Not panicking about alternatives. Even at double the price, Flash is the value king. The alternatives (GPT-5.6 Sol, Claude Fable 5) cost 60-100x more per test.
  4. Watching the free quota news. The 5T/day free tier on OpenCode is the part that actually matters for hobbyists like me. If the hike touches free quotas, that's the bigger story than any per-token price.
The bottom line

DeepSeek spent a year proving the "price war" wasn't a marketing phrase — then proved that being 100x cheaper than everyone has real consequences when everyone actually shows up. Eight trillion tokens a day isn't a growth story anymore; it's a cost story.

The fun part? Even with a "significant" increase, V4 Flash will almost certainly remain the cheapest serious model on earth. The price war's first casualty isn't a model — it's the price war itself.

Related on this blog: DeepSeek's Price Hike Gets Real: Up to 4.7x on Output · DeepSeek Added Vision to Flash. The Text Scores Went Up.

Numbers as of August 6, 2026. The exact new rates haven't been published yet — this post will need a follow-up the moment they drop.

Comments