Anthropic measures open model GLM-5.3's cyber offence skills
Anthropic says Zhipu's GLM-5.3 nearly matches Claude Mythos Preview at writing exploits, and its safeguards fail 64-100% of the time.
The test
On Chrome V8 vulnerabilities GLM-5.3 built working exploit chains in 50 of 410 attempts, against 56 for Claude Mythos Preview. A proof of concept chaining a new Chrome bug with a known flaw took 20 minutes of human time and eight hours of model time, costing $20.40 through Zhipu's API.
Safeguards
Framed as a red-team exercise, the model attempted to connect to the target in 64% of runs; prefilling its reasoning raised that to 92%, and removing refusals to 100%. The same attacks failed on safeguarded Claude models. Anthropic says governments should test capable models and defenders need tools as good as attackers have.
