Databricks' Yuchen Jin used a single price comparison to make the case that "intelligence per dollar" is the metric that shows how fast AI is moving: o1 Pro at $150 / $600 per million tokens 1.5 years ago, GLM-5.3-Flash at $0.15 / $0.50 today. Both numbers hold up against the vendors' own pages, which makes the headline arithmetic — a roughly 1,000x collapse — real rather than rhetorical.
OpenAI's o1-pro-2025-03-19 is still listed on the API at $150 per million input tokens and $600 per million output tokens, with a 200K context window, a 100K output cap and availability restricted to the Responses API. Z.ai's price page puts GLM-5.3-Flash's list rate at $0.15 input / $0.50 output, with cached input at $0.03.
| o1 Pro (March 2025) | GLM-5.3-Flash (today) | Ratio | |
|---|---|---|---|
| Input / M tokens | $150.00 | $0.15 | 1,000x |
| Output / M tokens | $600.00 | $0.50 | 1,200x |
| Context window | 200K | 1M | 5x |
The gap is momentarily wider than Jin's post says. GLM-5.3-Flash is running a 50% launch discount — $0.075 / $0.25 — that expires at 24:00 on September 9, 2026 Singapore time, which puts the live spread at 2,000x on input and 2,400x on output until the promotion lapses.
Jin's second assertion, that GLM-5.3-Flash is "more intelligent than o1 Pro," is the harder one, because almost nobody re-runs current evals against an 18-month-old reasoning model. Artificial Analysis's head-to-head page is the closest thing to a direct answer: it scores GLM-5.3-Flash at 42 on its current Intelligence Index against 12 for o1-pro, with blended prices of $0.098 per million tokens versus $195, and time-to-first-token of 1.91s against 29.27s.

GLM-5.3-Flash at the discounted rate sits on the global cost-versus-intelligence frontier, roughly an order of magnitude cheaper per task than the models scoring near it. Credit: chart published by Z.ai, sourced to Artificial Analysis.
On the newer v4.1.1 index used in the chart Z.ai publishes, GLM-5.3-Flash scores 57 points at $0.045 per task — three points below GLM-5.3 (max) and within a few points of Grok 4.6, GPT-5.6 Sol and Kimi K3, all of which cost between roughly $0.5 and $3 per task. (The 42 and the 57 are different index versions, not a contradiction.)
Z.ai's own benchmark numbers put the model against contemporaries rather than against o1 Pro: 84.3 on Terminal Bench 2.1, 63.4 on DeepSWE v1.1, 48.8 on AutomationBench, 55.3 on HLE with tools.

Z.ai's published results for GLM-5.3-Flash across six benchmarks. Credit: Z.ai.
The comparison is not quite like-for-like, and the ways it isn't are the interesting part. o1 Pro billed reasoning tokens as output at $600/M, so real task costs ran far above the sticker; GLM-5.3-Flash is an open-weights 320B-parameter mixture with 18B active, sold at a price a hosted lab could not match with a dense frontier model. Part of the 1,000x is a genuine efficiency gain — sparse plus linear attention cuts GLM-5.3-Flash's attention compute 3.01x and its KV cache 4.44x versus GLM-5.3 — and part of it is a Chinese lab pricing an open model as a land grab. Neither part is fake, but they decay differently: efficiency persists, promotional pricing does not.
For anyone budgeting an agent that burns tokens by the billion, the practical read is that the cost floor for near-frontier reasoning has dropped by three orders of magnitude in eighteen months, and the models sitting on today's frontier are the ones nobody was serving at any price when o1 Pro shipped.
Yuchen Jin's postZ.ai pricingZ.ai GLM-5.3-Flash overviewOpenAI o1-pro model pageArtificial Analysis: GLM-5.3-Flash vs o1-pro
Z.ai says it served GLM-5.3-Flash, OpenRouter's anonymous ox-alpha, entirely on Chinese chips
An 11x price spread for the same open-weight model on OpenRouter

OpenRouter's GPT-5.6 discount: Luna tokens up 13.8x, Terra 5.6x

Z.ai's GLM-5.3-Flash: two cheap attentions, MIT weights, and a claim it all ran on Chinese chips

OpenAI cuts GPT-5.6 Luna price by 80%

SemiAnalysis estimates different token-value ceilings for $200 AI plans