Back
SiTech Professional Insights
Z.ai releases GLM-5.3: stronger coding and emergent cyber capability
SiTech Team2 წთ. საკითხავი

Z.ai releases GLM-5.3: stronger coding and emergent cyber capability

Z.ai says GLM-5.3 keeps the same base model as GLM-5.2 and gets all its gains from scaled post-training: a 50 percent coding jump on its internal benchmark, top open-weights scores, and cyber capability that grew faster than expected.

Z.ai has released GLM-5.3, a new flagship model whose entire improvement over GLM-5.2 comes from post-training — the base model is unchanged. In its announcement on August 14, 2026, the company says the update delivers much better complex coding and long-horizon task performance, together with a cyber capability that developed faster than the team expected.

Coding: open-weights leadership, in Z.ai's benchmarks

Z.ai describes GLM-5.3 as the most capable open-weights model for coding, with a 50 percent improvement over GLM-5.2 on its in-house Z.ai Code Bench. The company reports best open-source results on Terminal Bench 3.0, where GLM-5.3 scores 28.3 against GLM-5.2's 4.6, and on Agents' Last Exam, 28.5 versus 23.8. On DeepSWE v1.1 it rises from 46.2 to 66.9. The margin to closed frontier models is not closed everywhere: on Terminal Bench 3.0, Fable 5 scores 33.7 and GPT-5.6 Sol 34.6.

The company also highlights token efficiency. At its maximum reasoning effort, GLM-5.3 reaches 34.5 percent on Z.ai Code Bench at roughly 75,000 output tokens per task, compared with 23.4 percent at 96,000 tokens for GLM-5.2.

Emergent cyber capability

Vulnerability-discovery data was added to the post-training mix in the hope of improving flaw analysis. According to Z.ai, the capability scaled further than expected: the model moved from spotting isolated bugs to planning multi-stage exploitation chains. It scores 84.5 percent on CyberGym, the best result reported on the benchmark, ahead of Mythos 5 at 83.8 and GPT-5.6 Sol at 83.6. On ExploitBench it more than doubles GLM-5.2, from 24.4 to 54.4, though Mythos 5 (78.0) and GPT-5.6 Sol (76.5) remain far ahead. Z.ai notes a clear pattern: the further up the exploitation chain a benchmark sits, the larger the gain — and the wider the remaining gap to closed models.

Real-world findings, and the release plan

Working with security teams in China, the company says its models have identified 2,436 vulnerabilities across 269 open-source projects after expert review and deduplication, including 1,097 rated critical or high. The oldest flaw dates to 1981, and the average vulnerability had lived 26.6 years before discovery. A public Security Disclosure Ledger tracks the findings: 53 are disclosed so far, with 2,383 still under embargo. The weights will be released about two weeks after launch, once safety evaluation and hardening are complete.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.