Anthropic says Z.ai's open-weight GLM-5.3 builds working cyber exploits almost as often as its own restricted Mythos model
50 exploits in 410 attempts against 56 for Claude Mythos Preview. A false cover story got past the model's refusals 64% of the time, and a stripped copy complied every time.
Anthropic published a research post on 29 September assessing GLM-5.3, the open-weight model from Z.ai, the company formerly known as Zhipu. Its conclusion is blunt: “The release of GLM-5.3 is a meaningful step change in the cyber capabilities available to attackers.”
On ExploitBench, GLM-5.3 built working end-to-end exploits in 50 of 410 attempts. Anthropic's own Claude Mythos Preview, which it does not release generally, managed 56 of 410. A model anyone can download now comes close to one Anthropic holds back.
Anthropic also tested the model's refusals. A plain request for an exploit was refused every time. A false cover story got through 64% of the time and prefilled reasoning 92% of the time, and an “abliterated” copy, with its refusals stripped out, complied 100% of the time. Because the weights are public, those refusals are not a control anyone can rely on. In a separate test, the smaller GLM-5.3-Flash exploited a known vulnerability, CVE-2026-11645, for $20.40, with 20 minutes of human attention and 8 hours of model time.
The post cites the US Center for AI Standards and Innovation (CAISI), which in a 17 September assessment called GLM-5.3 “the most cyber-capable open-weight model released to date” and placed it about four months behind the US frontier. Anthropic's ask: “Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3.” These are a competitor's measurements, and Z.ai's response is not visible from here.
- Confirmed On ExploitBench, GLM-5.3 built end-to-end exploits in 50 of 410 attempts, against 56 of 410 for Claude Mythos Preview. Anthropic
- Claimed Refusal bypass rates in Anthropic's tests: plain request 0%, false cover story 64%, prefilled reasoning 92%, abliterated copy 100%. Anthropic
- Claimed GLM-5.3-Flash exploited CVE-2026-11645 for $20.40, with 20 minutes of human attention and 8 hours of model time. Anthropic
Safety, security & governanceModels & releases
Today in the September 30, 2026 edition · front page