The Maivia Gazette

Verified AI news, every morning

Research

Anthropic says Zhipu's GLM-5.3 can build cyber exploits on its own and that its safeguards were bypassed in 64 to 100 percent of tests

The Frontier Red Team says exploit-building capabilities like Claude Mythos Preview's are now available in a model released without meaningful misuse protections.

Brass skeleton keys spill from an open toolbox across a workbench under a bright lamp.
AI-generated illustration, not event photography. The motion is AI-generated from the still.

Anthropic's Frontier Red Team has published an analysis of GLM-5.3, the latest model from Zhipu AI, which is known outside China as Z.ai. The team finds that, like Anthropic's own Claude Mythos Preview, GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits. Anthropic says the difference is that GLM-5.3 was released without meaningful safeguards to limit misuse. In the company's simulated tests, simple techniques bypassed the model's safeguards between 64 and 100 percent of the time, while the same attacks did not succeed against safeguarded Claude models. Anthropic explains that it released Mythos Preview five months ago only in a limited way, through Project Glasswing. That gave trusted defenders time to find more than 10,000 vulnerabilities in critical software before comparable models reached attackers. The company says that head start is now over, because "those models have now arrived." It assesses that GLM-5.3's lax safeguards make it significantly easier for malicious actors to carry out impactful cyberattacks. The findings come from a competitor's own testing and have not been independently verified. Even so, they suggest that defenders, software maintainers and policymakers may now face widely available exploit-building tools.

Sources

  1. AnthropicGLM-5.3 and the spread of advanced cyber capabilitiesPublished · fetched

Also in this edition