Skip to main content

Ox Alpha Was GLM-5.3-Flash, Trained on Chinese Chips Alone

Ox Alpha Was GLM-5.3-Flash, Trained on Chinese Chips Alone

August 26, 2026 · AI · LLM · Open Weights · GLM · Model Reviews

The anonymous model I wrote about three days ago finally has a name. Z.ai confirmed today that Ox Alpha — the free, week-long, nobody-will-claim-it frontier model on OpenCode — is GLM-5.3-Flash: a 320B-parameter, 18B-active multimodal MoE under MIT license. That part ends a guessing game. The part nobody's putting in the headline is the reveal buried in the announcement: it runs entirely on Chinese AI chips.

This is the follow-up I promised when I played detective on Ox Alpha, and honestly the outcome is more interesting than a name reveal. The mystery model wasn't just claimed by Z.ai — it's the first public look at what a frontier-class Chinese open-weights model costs when the training stack has zero NVIDIA hardware in it.

The reveal, finally

ItemGLM-5.3-Flash (Ox Alpha)
Parameters320B total, 18B active (A18B MoE)
LicenseMIT
ModalityNative multimodal (text + image + video)
Context1M tokens
Hardware"Runs entirely on Chinese AI chips"
Previous identityPreviewed anonymously as Ox Alpha

GLM-5.3-Flash, from Z.ai's August 26 announcement. MIT license is notable — a frontier-class model you can actually take.

I had flagged the MiMo-Vs-GLM question in my first post, with the tokenizer and video-encoder forensics pointing at Z.ai. The confirmation closes a loop that's become a genuine pattern: this is now the fifth anonymous model in about six months that turned out to be a Chinese lab, and the second from Z.ai after GLM-5 was quietly previewed as "Pony Alpha." The stealth-launch playbook is basically a marketing channel now.

Why "Chinese chips" is the real headline

Here's what's easy to skim past. Z.ai — formerly Zhipu — has been on the U.S. Commerce Department's entity list since January 2025, which cuts off legal access to NVIDIA silicon entirely. So the model you could use for free this week, the one benchmarking near GPT-5.6 mid, was not just built by a Chinese lab — it was built on a stack where the only accelerators available were domestic.

The numbers behind that are striking. GLM-5.2, Z.ai's previous release, was reportedly trained entirely on Huawei Ascend accelerators with no NVIDIA hardware involved, and it topped the open-weight leaderboards within a week of shipping. Then in July, Bloomberg reported Z.ai completed a 1-gigawatt data center built to run only on Chinese-made chips, with multiple clusters of more than 10,000 domestic accelerators each — and, critically, none of them NVIDIA. That's the infrastructure that trained the model you now know as GLM-5.3-Flash.

This matters because the conventional wisdom has been that Chinese labs are lagging on chips, that export controls would hold them back. And yet here's an open-weights model, free for a week, scoring against the US frontier on a substack Z.ai built without a single NVIDIA part. The chips — likely Huawei Ascend, on the training history — still trail Blackwell on raw performance-per-watt, so a gigawatt of Chinese silicon delivers less usable compute than a gigawatt of NVIDIA systems. But Z.ai is proving that gap no longer makes a Chinese frontier-class model impossible.

The performance, redrawn from their own chart

Bar charts: GLM-5.3-Flash vs GLM-5-2, DeepSeek, Claude Opus 4.8, GPT-5.6 Turbo, Gemini 3.7 Flash across six benchmarks

Z.ai's own launch comparison, redrawn. Vendor-reported — and it's honestly a mixed bag, which is the useful part. Chart by the author.

This is where I want to be careful, because the chart tells a more honest story than "it wins everything." On Terminal Bench 2.1, GLM-5.3-Flash scores 84.3 — above its own GLM-5-2 (81.0) and DeepSeek (83.9), but below GPT-5.6 Turbo (87.6) and Gemini 3.7 Flash (86.8). On DeepSWE-v1 it actually sits at the bottom of the field (43.4), behind the whole group including Claude, DeepSeek, and even GLM-5-2. It wins Code w/ Tools (68.5, ahead of the 66.3 the Flash line's ancestor managed), but the agentic-coding lead is not the slam dunk the "frontier" framing implies.

So the honest read is: this is a very solid mid-to-upper open-weights coding model at a 320B/18B size, priced aggressively and fully open under MIT — but the "beats GPT-5.6 mid" tallies come from independent spot measurements (like the ~63% 113-task DeepSWE run the community did on Ox Alpha), not from a clean victory lap in Z.ai's own six-benchmark table. I'll believe the margins once an independent leaderboard lands.

MIT matters more than the benchmarks

The license is the quietly biggest part of this release. A 320B multimodal model with vision and video, an 18B active count, native 1M context — and it's MIT. No Qwen-style Community License counting your revenue, no "AI Work Assistant" carve-out. You can take it, fine-tune it, sell a hosted service on it, and not ask anyone's permission. That's the model Z.ai is betting on to win the same game the stealth release just did: get it into hands, into stacks, into production, and win on being the open option.

My take

I asked in the first post whether Ox Alpha's identity even mattered. The answer, it turns out, was yes — but not for the reason I thought. The reveal that it's GLM-5.3-Flash is less interesting than the reveal that it's Chinese-chip-only. A week after sitting on the OpenCode mystery board, we now know what an export-restricted lab can ship when it has no choice but to build its entire stack on domestic silicon: a genuinely competitive, fully open, multimodal 320B model. The chips still cost more compute per watt, and I'm not about to pretend the ASCEND stack matches Blackwell efficiency — but the ceiling Z.ai just demonstrated is higher than the export-control narrative has assumed.

The MIT license is the part I'd build on. The benchmarks are a mixed bag and worth waiting for independent numbers on; the open access is not. Watch for whether the anonymous-release playbook keeps producing winners like this, and whether US labs start answering the "what can a no-NVIDIA lab ship" question — because it just got a very confident answer.

Sources: Z.ai official announcement (Aug 26 2026), the X post announcing GLM-5.3-Flash and its benchmark chart, previous coverage in this blog (Ox Alpha, GLM-5.3), Tom's Hardware / Bloomberg / The Next Web on Z.ai's Chinese-chip data center, CommunitySpot's independent DeepSWE measurement of Ox Alpha. All benchmark figures in the redrawn chart are Z.ai's vendor-reported values from their announcement; the 63% DeepSWE figure is an independent community measurement. GLM-5.2's Ascend-only training and the 1GW data-center details are press-reported, not confirmed by Z.ai.

Comments