Back
GLM-5.2 becomes the leading open weights model on Artificial Analysis
SiTech AI Team2 წთ. საკითხავი

GLM-5.2 becomes the leading open weights model on Artificial Analysis

Z.ai's GLM-5.2 scores 51 on the Artificial Analysis Intelligence Index v4.1, ahead of MiniMax-M3, DeepSeek V4 Pro and Kimi K2.6, and sits on the Pareto frontier of intelligence versus cost per task.

Z.ai's GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index, scoring 51 on version 4.1 of the benchmark. According to the analysis firm, the model keeps the same size as its predecessor GLM-5.1 — 744 billion total parameters with 40 billion active — while gaining 11 points on the index.

Where it stands

The score puts GLM-5.2 ahead of other open weights releases: MiniMax-M3 and DeepSeek V4 Pro (max) both score 44, while Kimi K2.6 reaches 43. On GDPval-AA v2, Artificial Analysis's primary metric for real-world agentic performance, GLM-5.2 scores 1,524, ahead of MiniMax-M3 (1,418) and DeepSeek V4 Pro (max, 1,328), and effectively level with GPT-5.5 in its xhigh reasoning configuration (1,514).

Gains across evaluations

The improvement over GLM-5.1 shows up in most evaluations, with scientific reasoning leading the way. CritPt rises 16 points to 21 percent, HLE gains 12 points to 40 percent, and SciCode adds 7 points to reach 50 percent. AA-LCR climbs 9 points to 71 percent, the tau3 banking benchmark gains 15 points to 27 percent, and TerminalBench v2.1 improves 16 points to 78 percent. GPQA Diamond edges up 3 points to 89 percent.

Cost and reliability

GLM-5.2 also sits on the Pareto frontier of the intelligence versus cost per task chart, meaning it has the lowest cost per task among models at its intelligence level. On the AA-Omniscience Index it scores 4, up from 2 for GLM-5.1, with accuracy rising to 25.1 percent from 24.2 percent and the hallucination rate falling to 28.1 percent from 29.4 percent, while the attempt rate stays flat at 47 percent. The model uses about 43,000 output tokens per Intelligence Index task, of which 37,000 are reasoning tokens.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.