GLM-5.2 was too expensive for its own hype. Then it went 95% off.
GLM-5.2 was too expensive for its own hype. Then it went 95% off.
When GLM-5.2 launched in June, the internet had two reactions in quick succession. First: "wait, this open-weight model is beating Claude Opus in agent benchmarks?" Then, about thirty seconds later: "okay but who is paying that API price." Two months later, OpenRouter is selling the same model at 95% off — and I genuinely think this is the moment the hype and the price finally meet.
June: the hype was real
GLM-5.2 came out June 16 as an open-weight (MIT) reasoning model with a 1M-token context window — 753B parameters, 40B active. The reception was not lukewarm. Hacker News had a thread literally titled "GLM 5.2 beats Claude in our benchmarks", independent testers found it winning four of five practical building tasks against Claude Opus 4.8, and the model card numbers backed the noise up:
For a model you could self-host, with weights you could actually download, that was a statement. The community was genuinely excited — this felt like the point where open-weight caught up to the closed frontier.
And then everyone looked at the price
Z.ai's official API rate was $1.40 per million input tokens, $4.40 output. Not absurd by frontier standards — but this was an open-weight model competing in a market where the cheap tier had trained everyone to expect single-digit prices. The reaction split cleanly down the middle: subscription users got the $18-$65/month GLM Coding Plan and mostly shrugged happily ("Z.ai doesn't use token-based pricing in its subscription system, which works to my advantage lol" — actual Reddit quote), while API users grumbled. One HN commenter on the $65 Pro tier said the weekly limits ran out "somewhat fast" and that quota multipliers (3x during peak hours) were getting worse after September.
The verdict that stuck was a Reddit thread title: "GLM 5.2 looks strong — but does the price gap still make sense?" Strong model, wrong price for what it was. That's where the story sat for two months.
August: the price war nobody announced
OpenRouter doesn't set these discounts — the providers on it do. And sometime in the last few weeks, three hosting companies decided GLM-5.2 traffic was worth losing money on:
- StreamLake — 95% off: $0.07 in / $0.22 out (official: $1.40 / $4.40)
- NovitaAI — 94% off: $0.08 in / $0.26 out
- Baidu Qianfan — 92% off: $0.11 in / $0.35 out
- Decart — 40% off: $0.72 / $1.80, for contrast
Cache reads got cut the same way — $0.013 per million at StreamLake, 95% off — which is the number that matters for agent loops that re-send context every turn. The June complaint was "great model, meh price." The August question is: does this fix it?
Does the discount make GLM-5.2 competitive again?
Let's put a real workload on it: 100 million input tokens a month, 20 million output, half the input hitting cache — a normal agent-session pattern:
| Route | Monthly cost |
|---|---|
| OpenRouter via StreamLake (95% off) | ~$8.55 |
| OpenRouter via NovitaAI (94% off) | ~$9.92 |
| OpenRouter via Baidu (92% off) | ~$13.68 |
| OpenCode Go subscription — GLM-5.2 at full rate | $10 flat, capped at $60 usage/mo (~35% of this workload) |
| Z.ai official API (no discount) | ~$171 |
Here's the honest answer to the question. For API users: yes, it's competitive again — arguably the best value on the market. A frontier-class reasoning model at $0.07/$0.22 per million, with discounted cache, lands below the price floor the cheap tier had settled on. The June story was "great model, wrong price." The discount deletes exactly that sentence. The hype and the price have finally met.
One row deserves extra explanation: OpenCode Go's $10 flat sounds like the cheapest row, but GLM-5.2 there is billed at the full official rate inside a $60/month usage cap (about 4,300 requests). My workload — 50M fresh + 50M cached input, 20M output — comes to ~$171 of usage at those rates, so the subscription covers roughly a third of it before you hit the wall. The 95%-off price exists only on OpenRouter's discounted providers. Same for the Coding Plan: Z.ai's own subscription doesn't get cheaper either.
But two asterisks. First, the discount lives on OpenRouter's providers, not on Z.ai's own rate card — the official API still charges $1.40/$4.40, so this helps API users who route through OpenRouter. Second, this is a provider price war, not a Z.ai announcement: there's no end date, and there's also no guarantee. More on that below.
Before you celebrate: four caveats
1. The discount is a provider decision, not a law of nature. StreamLake, NovitaAI and Baidu are fighting for OpenRouter traffic — that's the whole reason this exists. No end date on the page, but no contract either. Provider prices change without notice; that's how this market has always worked. I'd bet on the war lasting (three competitors, one model, 20x spread), but "bet" is the operative word.
2. Reasoning tokens balloon. GLM-5.2 is a reasoning model and 'high'/'xhigh' effort burns output tokens fast. At $0.22/M output that's still cheap, but your bill can double or triple versus what input-price math suggests. Watch actual output volume, not just the rate.
3. The cache discount is OpenRouter-specific. Z.ai's own API charges full price on cached input. The 95%-off cache read exists only on these providers — route elsewhere and the economics change.
4. The Coding Plan crowd gets nothing from this. Subscribers already had their flat-rate deal; this discount is aimed at API traffic. If you're on the $18-65/month plan, your calculus doesn't move.
So: does it have competitiveness back?
For the API crowd — the people who spent June saying "great model, wrong price" — yes. This is the first time GLM-5.2's actual capability and its actual price have been in the same league. I switched my heavy agent sessions over, and at these numbers the honest answer to "should I try it" is just: yes, obviously, it's eight dollars a month. The model that was too expensive for its own hype is now priced like the hype was right all along.
I'll report back in a few weeks on whether it holds — if the war ends or the reasoning tokens eat the savings, you'll hear it here first. Meanwhile, if you're paying the official API rate: stop. Log into OpenRouter. The June problem has a two-month-old solution now.
Prices checked Aug 10, 2026. Launch reactions: Hacker News (item 48709670, 48714425), r/artificial, r/AISEOInsider, r/ZaiGLM. Model card scores: Z.ai official. OpenCode Go limits: opencode.ai/docs/go. Costs assume 100M input / 20M output / 50% cache-hit per month at listed rates.
Related on this blog: Astra: 10 Math Proofs, Then OpenAI Paused Its Own Model · 96% to 11%: GLM-5.3-Flash, Uncensored at the Weight Level
Comments
Post a Comment