Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

anthropicanthropicanthropicAnthropic published an analysis on Tuesday warning that GLM-5.3, the latest AI model from Chinese developer Zhipu AI (also known as Z.ai), can autonomously build end-to-end cyber exploits on par with Anthropic's own Claude Mythos Preview — but has been released as an open-weight model without effective safeguards against misuse.anthropic
The report from Anthropic's Frontier Red Team found that GLM-5.3 developed working exploits in 50 of 410 attempts on ExploitBench, a benchmark measuring exploit development against known Chrome Alphabet Inc. V8 engine vulnerabilities, compared with 56 for Claude Mythos Preview. On a binary exploitation benchmark, GLM-5.3 achieved full control-flow hijacks in 4 percent of trials versus 6 percent for Mythos. Earlier models, including Claude Opus 4.6 and GLM-5.2, succeeded in none.anthropic
Anthropic's central concern is not capability alone but the absence of guardrails. Because GLM-5.3's weights are publicly available, its built-in refusals can be removed through a technique called "abliteration." Anthropic's team — with no prior experience with the method — reduced the model's refusal rate from above 90 percent to roughly 3 percent on standard harm benchmarks at a cost of about $4,400 in compute. Several developers published similarly unlocked versions within days of the model's August release.anthropic
Short of full abliteration, simpler tricks also work. A deceptive prompt claiming to be an authorized red-team exercise bypassed GLM-5.3's safeguards 64 percent of the time; prefilling the model's chain-of-thought tokens raised compliance to 92 percent. None of these techniques succeeded against Claude models in Anthropic's testing.anthropic
In one researcher-driven session, GLM-5.3 discovered several previously unknown vulnerabilities in a popular web browser's JavaScript engine and chained them into a working exploit — a webpage that reads arbitrary files from a visitor's computer. The smaller GLM-5.3-Flash built a reliable exploit chain for a known Chrome vulnerability in eight hours of compute and 20 minutes of human attention, at an API cost of $20.40.anthropic
On Sept. 17, NIST's Center for AI Standards and Innovation published its own evaluation, calling GLM-5.3 "the most cyber-capable open-weight model released to date" while noting its capabilities remain about four months behind current U.S. frontier models. Z.ai first launched GLM-5.3 through its API in August and released the weights on Hugging Face two weeks later, saying the delay allowed for additional safety evaluations.trendingtopics+2
Not everyone shares Anthropic's alarm. Jake Williams of IANS Research told The New Stack that while threat actors will use the model, he does not expect it to meaningfully shift the threat landscape. Critics have also pointed to a potential conflict of interest, noting that Anthropic competes directly with Z.ai and was the only major AI lab not to sign an industry letter supporting open models backed by Nvidia , Meta , Microsoft , OpenAI, and Google.trendingtopics
Anthropic framed the moment as one that demands expanded access for defenders: "A critical threshold in freely accessible capabilities has now been crossed".anthropic