Skip to main content
AI Socratic
← News
z.AI

Company / organization

z.AI

Website / profile ↗

By and about z.AI

news

From o1 Pro to GLM-5.3-Flash: a 1,000x fall in token prices in 18 months

Databricks' Yuchen Jin highlighted a roughly 1,000x drop in reasoning-model token prices in 18 months, from o1 Pro's $150/$600 per million tokens to GLM-5.3-Flash's $0.15/$0.50 today (or $0.075/$0.25 during its launch discount).

news

Z.ai releases GLM-5.3 weights

Z.ai announced on August 28 that GLM-5.3 is now open-weight, released via a single post from its @Zai_org account with no accompanying license terms or benchmark tables.

news

Z.ai's GLM-5.3-Flash: two cheap attentions, MIT weights, and a claim it all ran on Chinese chips

Z.ai revealed OpenRouter's anonymous ox-alpha as GLM-5.3-Flash, a 320B-A18B MIT-licensed MoE that pairs sparse and linear attention to cut KV cache 4.44x, scores 57 on Artificial Analysis's index, and it says was served entirely on Chinese chips.

news

Z.ai says it served GLM-5.3-Flash, OpenRouter's anonymous ox-alpha, entirely on Chinese chips

Z.ai revealed that ox-alpha, the anonymous model that led OpenRouter for six days with 23.2T tokens, is GLM-5.3-Flash, and says it was served entirely on 100,000+ domestic Chinese chips at $0.15/$0.50 per million tokens.

news

Chinese labs converge on one architecture: 3:1 linear attention and a 2,048-token budget

Z.ai's GLM-5.3-Flash and Alibaba's Qwen3.8-Flash-Next, released a day apart, independently converged on the same recipe: 3:1 linear attention, a 2,048-token attention budget and four-branch gated residuals, while MiniMax dissents and keeps full attention.

news

MazeBench: the best coding agent collects 13 gems, four frontier models get zero

MazeBench's code-enabled leaderboard has GPT-5.6 Sol topping the board with 13 of 100 hidden gems, claude-opus-5 at 12% and claude-fable-5 at 11%, while grok-4.6, ox-alpha, glm-5.3 and qwen3.8-max all scored 0%.

news

GLM-5.2: Z.ai’s Open-Weight Beast for Long-Horizon Coding

Z.ai released GLM-5.2, a 744B open-weight model scoring 62.1% on SWE-bench Pro and rivaling closed models like Claude Opus on coding tasks, with 1M token context and roughly one-sixth the cost of GPT-5.5.

blog

AI Socratic June 2026 #2 — Begun the Open Source AI War Has

The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.

news

GLM 5.2

Z.ai released GLM 5.2, an open-weights reasoning model with a 1M-token context window designed for long-horizon agent workflows, software engineering, and multi-step automation tasks.

news

GLM 5.1

Z.ai released GLM-5.1 on April 3, 2026, an open-weights model with major gains in coding capability and long-horizon task handling.

news

GLM 4.7 Flash

Zhipu AI released GLM-4.7-Flash, a 30B open-weights model positioned as a performance-efficiency balance in the SOTA class.

news

GLM 4.6

Z.ai released GLM-4.6 on September 29, 2025, expanding the context window from 128K to 200K tokens and improving coding performance, reasoning, and tool use capabilities.