
Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits
Anthropic's Frontier Red Team says Zhipu's open-weight GLM-5.3 can build complete cyber exploits on its own, and that its safeguards can be bypassed between 64% and 100% of the time with simple techniques. The model shipped without meaningful restrictions.
Anthropic's Frontier Red Team published an analysis on September 29 of GLM-5.3, the open-weight model from China's Zhipu AI. The company says the model can build complete, end-to-end cyber exploits on its own, much like Anthropic's own Claude Mythos Preview, which arrived five months earlier with limited access under Project Glasswing. Zhipu AI operates in China and is known as Z.ai elsewhere. The key difference is that GLM-5.3 shipped with open weights and no meaningful safeguards.
What the tests show
On ExploitBench, which measures how well models exploit known flaws in Google Chrome's V8 engine, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts, reaching 12%, while Claude Mythos Preview reached 14%. In Anthropic's internal Binary Exploitation benchmark, GLM-5.3 achieved a full control-flow hijack in 4% of trials and Mythos Preview in 6%. Earlier models, including Claude Opus 4.6 and GLM-5.2, succeeded in none of them. NIST's Center for AI Standards and Innovation (CAISI) called GLM-5.3, in a September 17 assessment, the most cyber-capable open-weight model released to date, lagging the US frontier by about four months.
Bypassing the safeguards
GLM-5.3 does ship with built-in safeguards and often refuses clearly harmful requests, but in Anthropic's tests those safeguards were bypassed between 64% and 100% of the time using simple techniques. A deceptive cover story, telling the model it is an autonomous red-team agent on an exercise, raised engagement to 64%; prefilling its reasoning tokens raised it to 92%; and abliteration, which is possible because the weights are public, raised it to 100%.
Anthropic produced its own abliterated copy, a task that took its team about 2,200 GPU hours and roughly $4,400. The edit cut the model's refusal rate from above 90% to about 3% and 2% on JailbreakBench and HarmBench, and to 12% on StrongREJECT. Abliterating GLM-5.3-Flash took about 600 GPU hours. None of these techniques worked against safeguarded Claude models in the testing.
What it means
In one session, a researcher used GLM-5.3-Flash to build an exploit for a recently disclosed Chrome flaw, CVE-2026-11645. With little direction, the model chained two exploits into a reliable chain for an ARM64 target, bypassing pointer-authentication hardening. It took 20 minutes of human attention plus eight hours of model work, costing $20.40 at Zhipu's API prices. In another session the model found chains of 0-day flaws in a browser component and read a user's SSH private key from their computer through a malicious website.
Anthropic argues that defenders should not be left with models weaker than their adversaries, and calls on governments to conduct independent safety testing of sufficiently capable models.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.