Anthropic Finds Downloadable GLM-5.3 Nears Claude’s Exploit-Building Ability
The benchmark results put Zhipu’s model close to Claude Mythos Preview. Separate simulations tested willingness to attack, not whether attacks would succeed.
Loading page…
The benchmark results put Zhipu’s model close to Claude Mythos Preview. Separate simulations tested willingness to attack, not whether attacks would succeed.
Listen to this story
Anthropic’s September 29 tests put GLM-5.3 close to Claude Mythos Preview on two measures of building working exploits: 50 versus 56 completed V8 exploits, and 4% versus 6% execution takeovers. In guided tests, GLM-5.3 chained browser flaws into an exploit that read files, while GLM-5.3-Flash combined known flaws with an estimated $20.40 in model-use costs. Simulated bypasses also weakened GLM-5.3’s refusals. The evaluations used controlled targets and do not show attacks on real victims.
GLM-5.3 engaged with simulated malicious requests in 64% of runs after a deceptive authorization claim and 92% after reasoning was prefilled; engagement did not mean a successful compromise.
Removing refusal behavior dropped refusal rates from above 90% to 2%–12% across three harmful-request benchmarks, while scientific-test performance was unchanged.
NIST’s September 17 assessment ranked GLM-5.3 the leading open-weight model for cyber capability, about four months behind the U.S. frontier; some comparison models were restricted to vetted users.
An AI model anyone can download is approaching the exploit-building ability Anthropic previously put behind restricted access. In a September 29 evaluation, Anthropic found that Zhipu AI’s GLM-5.3 came close to Claude Mythos Preview on two cyber benchmarks—and that its refusal safeguards were readily bypassed in separate simulated tests.
Anthropic ran the exploit evaluations in isolated environments against offline targets set up for testing. These are controlled demonstrations of capability, not evidence that GLM-5.3 carried out the same attacks against real victims. The comparison also separates raw capability from deployed safeguards: the Claude benchmark runs had safeguards disabled.
ExploitBench tests whether models can turn known vulnerabilities in Chrome’s V8 JavaScript engine into working exploits. GLM-5.3 completed 50 of 410 attempts, versus 56 for Mythos Preview. The result concerns complete exploit development, rather than simply spotting a bug or describing a possible attack.
Anthropic’s internal binary-exploitation test uses open-source projects participating in Google’s OSS-Fuzz program. Full credit requires taking over a program’s execution. Across 100 randomly selected tasks, GLM-5.3 succeeded in 4% of trials and Mythos Preview in 6%. Earlier Claude Opus 4.6 and GLM-5.2 models had no successes on this test.
In a researcher-guided session, Anthropic says GLM-5.3 found previously unknown flaws in a browser’s JavaScript engine. Within a day, with limited human attention, it connected them into a webpage that could read arbitrary files from the tested Linux computer. Anthropic says it disclosed those vulnerabilities to the browser maintainer.
A second session tested the smaller GLM-5.3-Flash against two known vulnerabilities. It combined them into an ARM64 exploit that bypassed pointer authentication, a processor-level security protection. Anthropic reports eight hours of model work and 20 minutes of human attention. At Zhipu’s API prices, the estimated model-use cost was $20.40—not a price for developing every exploit.
GLM-5.3 initially refused every overtly malicious request in Anthropic’s simulated attack test. That changed under three bypass conditions:
Engagement meant trying to connect to a target, not successfully compromising it. The Decoder notes that this simulation did not execute code. Safeguarded Claude models showed zero engagement, although two bypass methods were unavailable against Claude: users cannot edit its model weights or prefill its reasoning through Anthropic’s API.
Removing refusals through a technique called abliteration cost Anthropic roughly $4,400 and 2,200 GPU hours. Refusal rates fell from above 90% to between 2% and 12% across three harmful-request benchmarks. Scientific-test performance was unchanged, while performance on a tested cyber subset fell only a few percentage points.
Anthropic also cites a September 17 assessment from NIST’s Center for AI Standards and Innovation. It ranked GLM-5.3 as the most cyber-capable open-weight model to date, roughly four months behind the U.S. frontier across its cyber benchmarks. That comparison included U.S. models available only to vetted users and tested with safeguards disabled where applicable.
Anthropic argues that the same capabilities can benefit defenders and calls for government safety testing of sufficiently capable models, including GLM-5.3’s successors. Its forecast that malicious actors will use such models to cause real-world harm remains a forecast, not an outcome demonstrated by these tests.
Editorial analysis
The strategic pressure here is on access policy, not just model rankings. If defenders can obtain comparable exploit-building ability outside a vetted program, restricted providers must explain what their controls preserve and whom they exclude. Anthropic’s evidence supports concern about removable refusals, but its safeguards comparison also favors its own closed-weight deployment model. Our view: independent evaluations should separate capability, willingness to comply, and actual attack success rather than fold them into one safety verdict. The next evidence to watch is independent testing of GLM-5.3’s successors, which Anthropic explicitly recommends—not another near-parity benchmark score alone.
Loading discussion...
Join the conversation
Explain whether user control or preventing misuse should carry more weight.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.