
Z.ai announced GLM-5.3-Flash on August 26, the latest speed-tier release in its GLM line. It is the model that had been running as Ox Alpha: a 320B-A18B mixture of experts under the MIT License, natively multimodal, with a 1M-token context window.

The launch card: 320B-A18B, MIT-licensed, 1M context, trained and served on Chinese AI chips — @Zai_org
The detail drawing the most attention is where it runs. SemiAnalysis puts the model at 100 trillion tokens per day served on Chinese chips — a claim about supply chains as much as about model quality.

SemiAnalysis on the serving footprint behind the release — @SemiAnalysis_
Artificial Analysis has already published its own intelligence, performance and price analysis of the model — the usual independent read on where a -class release actually lands on the cost/quality frontier, as opposed to where the launch post says it does. Their number: , on the Pareto frontier.

Artificial Analysis puts GLM-5.3-Flash at 57 intelligence, $0.09 per task — @ArtificialAnlys
Arena has it landing around #5 in Code Arena: WebDev — second among open models — at 1634 AutoEval, priced at $0.15/$0.5 per Mtoken.

Arena: ~#5 on Code Arena WebDev, #2 among open models — @arena
On the architecture, Sebastian Raschka reads GLM-5.3-Flash as a shift to a Kimi Linear-style 3:1 hybrid attention pattern — 34 Kimi Delta Attention layers against 11 MLA / DeepSeek Sparse layers — which is where the cheaper long-context serving comes from.

Sebastian Raschka on the hybrid attention layout versus GLM-5.2 — @rasbt
Worth pulling both up side by side before you wire it into anything: with fast, cheap tiers the interesting numbers are throughput and price per million tokens, and those are the ones vendor charts tend to present most generously.
A speed-tier model that reaches the frontier on cost, ships under MIT, and is served at scale on non-NVIDIA silicon changes three arguments at once — about price, about openness, and about whether export controls bind. OpenRouter, which ran it anonymously as Ox Alpha, reports over 20 trillion tokens in six days, the largest volume it has recorded for any model.

OpenRouter: Ox Alpha was its highest-volume model ever, 20T tokens in six days — @OpenRouter
Sources: Z.ai announcement · Artificial Analysis · OpenRouter model page · HN discussion · @Zai_org · @SemiAnalysis_ · @ArtificialAnlys · @arena · @rasbt · @Hesamation · @OpenRouter