GLM-5.3 Cyber Capabilities: Anthropic's Findings
Anthropic put GLM-5.3 cyber capabilities at 50 working exploits in 410 tries, against Claude's 56. The capability gap closed, the safeguard gap did not.
Anthropic published red-team results on 29 September 2026 putting GLM-5.3 cyber capabilities within a few points of its own frontier model. GLM-5.3, an open-weight release from Zhipu AI, built working software exploits in 50 of 410 attempts. Claude Mythos Preview managed 56 of 410.
That is a gap of six. The interesting finding is not the six. It is what sits on either side of the two models.
The GLM-5.3 cyber capabilities Anthropic measured
Measure | GLM-5.3 | Claude Mythos Preview |
|---|---|---|
ExploitBench successes | 50 of 410 | 56 of 410 |
Internal binary exploitation benchmark | 4% | 6% |
Safeguards bypassed, red-team framing | 64% of runs | not applicable, gated access |
Safeguards bypassed, prefilled reasoning | 92% of runs | not applicable |
Safeguards bypassed, after abliteration | 100% of runs | not applicable |
The Decoder's summary of the research records the bypass figures, and they are the story. GLM-5.3 does refuse malicious requests when asked plainly. Reframing the same request as a red-team exercise got compliance in 64 percent of runs. Prefilling the model's reasoning pushed that to 92 percent. Applying abliteration, a published technique for stripping refusal behaviour out of open weights, reached 100 percent.
Why open weights change the calculation
A capable model behind an API has a control surface. The provider can rate-limit, log, require identity, refuse a customer, or revoke access after the fact. Anthropic's argument is that Mythos Preview sits behind exactly that, and GLM-5.3 does not, because once weights are downloadable the safeguards are a suggestion rather than a boundary. Abliteration reaching 100 percent is the demonstration of that point, not an incidental detail.
This is also not Anthropic's claim alone. NIST's AI safety institute CAISI assessed GLM-5.3 in September and called it the most cyber-capable open-weight model to date, positioning it roughly four months behind the best US systems. Two organisations with different incentives measured the same thing and landed in the same place, which is worth more than either assessment on its own.
Reading a vendor's assessment of a competitor
Anthropic benchmarked a rival's model and concluded that rival's release practices are unsafe. The conflict of interest is obvious and should be stated rather than ignored. Three things make this report more credible than the usual version of that genre:
It published its own model's score alongside, and its own model scored higher on exploit generation. A pure hit piece does not hand you that number.
The bypass methodology is named and reproducible, not asserted. Abliteration is a public technique.
An independent government evaluator reached a comparable conclusion two weeks earlier, by a different route.
What it does not establish is real-world harm. An ExploitBench success is a benchmark artifact, not an attack in the wild, and a 12 percent success rate on exploit construction is a long way from a reliable capability. The honest reading is that the trend line crossed a threshold, not that something happened.
What this means if you build software
The practical consequence is not that you are now a target of AI-written exploits. It is that the cost of finding a vulnerability in your code dropped for everyone, including people who were never going to pay for API access. Anthropic's own recommendation is the defender's version of the same observation: governments should test capable models, and defenders need tools at least as good as the ones their adversaries have.
For the defensive side of this, whether AI can find security vulnerabilities in your code covers what the same capability does when pointed at your own repository, and how to catch an AI coding agent introducing a vulnerability covers the nearer-term risk for most teams, which is still your own agent rather than someone else's.
For what the model itself is, with a pricing and context-window comparison against its peers, see our GLM-5.3 explainer. The access-model question underneath this story is the subject of open-weight vs closed AI models. For the pattern of labs publishing capability thresholds, see OpenAI's critical cybersecurity threshold and the AI cyber defense open letter.
Gigazine's write-up carries the same figures with additional detail on Zhipu's release terms. Zhipu AI, which operates as Z.ai outside China, had not published a response to the Anthropic findings at the time of writing.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


