Z.ai Launches GLM-5.3-Flash: Frontier-Level AI at One-Tenth the Price

Chinese lab Z.ai released GLM-5.3-Flash, a natively multimodal model with 320B parameters and just 18B active. It beats GLM-5.2 at one-tenth the price and nears Claude Opus 4.8 on coding and agent benchmarks.
What happened
On August 26, Chinese AI lab Z.ai released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series. It carries 320 billion total parameters but only 18 billion active ones, beats its predecessor GLM-5.2 on benchmarks at roughly one-tenth the price, and gets close to Claude Opus 4.8 on coding and agentic tasks.
An architecture built for cheap inference
The model combines sparse and linear attention in a hybrid design and adds Manifold-Constrained Hyper-Connections (mHC) for more efficient scaling. Against the GLM-4.5 series, Z.ai nearly halved the number of layers (45 versus 92) and activation size, keeping serving costs low even at a one-million-token context.
Numbers and a quiet test run
Before launch, the model ran anonymously as "ox-alpha" on OpenCode and OpenRouter, quickly becoming the most popular model of the week — all traffic served on Chinese AI chips. It scores 57 on the Artificial Analysis Intelligence Index v4.1.1 at about $0.045 per task, and reaches 63.4 versus 46.2 for GLM-5.2 on DeepSWE v1.1.
Why it matters
For developers and agencies the frontier-to-budget gap keeps shrinking: top-tier coding and agent performance at a fraction of the cost changes what is realistic to automate, and price pressure on flagship models will only grow.